| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat: FSDP2 Longcat-Flash-Lite adaptation for the GEMM function Co-authored-by: wj<wangjin230@huawei.com> # message auto-generated for no-merge-commit merge: !4753 merge fsdp2-gemm into master feat: FSDP2 Longcat-Flash-Lite adaptation for the GEMM function Created-by: gcw_RxnYoBVv Commit-by: wj Merged-by: ascend-robot Description: ## What this PR does / why we need it? LongCat-Flash-Lite supports online weight conversion, Grouped GEMM (GMM) expert computation, and N-gram embedding memory optimization in the FSDP2 scenario. ## Does this PR introduce any user-facing change? The model can directly use the original HuggingFace weights to start FSDP2 training/inference, and reduce the risk of out-of-memory (OOM) caused by dense gradient accumulation in the backward phase of multiple N-gram embedding tables. ## How was this patch tested? Single-machine layer reduction precision alignment, with an absolute error of 0.004. See merge request: Ascend/MindSpeed-LLM!4753 | 14 天前 | |
feat: FSDP2 Longcat-Flash-Lite adaptation for the GEMM function Co-authored-by: wj<wangjin230@huawei.com> # message auto-generated for no-merge-commit merge: !4753 merge fsdp2-gemm into master feat: FSDP2 Longcat-Flash-Lite adaptation for the GEMM function Created-by: gcw_RxnYoBVv Commit-by: wj Merged-by: ascend-robot Description: ## What this PR does / why we need it? LongCat-Flash-Lite supports online weight conversion, Grouped GEMM (GMM) expert computation, and N-gram embedding memory optimization in the FSDP2 scenario. ## Does this PR introduce any user-facing change? The model can directly use the original HuggingFace weights to start FSDP2 training/inference, and reduce the risk of out-of-memory (OOM) caused by dense gradient accumulation in the backward phase of multiple N-gram embedding tables. ## How was this patch tested? Single-machine layer reduction precision alignment, with an absolute error of 0.004. See merge request: Ascend/MindSpeed-LLM!4753 | 14 天前 | |
feat: FSDP2 Longcat-Flash-Lite adaptation for the GEMM function Co-authored-by: wj<wangjin230@huawei.com> # message auto-generated for no-merge-commit merge: !4753 merge fsdp2-gemm into master feat: FSDP2 Longcat-Flash-Lite adaptation for the GEMM function Created-by: gcw_RxnYoBVv Commit-by: wj Merged-by: ascend-robot Description: ## What this PR does / why we need it? LongCat-Flash-Lite supports online weight conversion, Grouped GEMM (GMM) expert computation, and N-gram embedding memory optimization in the FSDP2 scenario. ## Does this PR introduce any user-facing change? The model can directly use the original HuggingFace weights to start FSDP2 training/inference, and reduce the risk of out-of-memory (OOM) caused by dense gradient accumulation in the backward phase of multiple N-gram embedding tables. ## How was this patch tested? Single-machine layer reduction precision alignment, with an absolute error of 0.004. See merge request: Ascend/MindSpeed-LLM!4753 | 14 天前 |