| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat:Add adaptation support for GLM5.2 model Co-authored-by: LinShua<707894133@qq.com> # message auto-generated for no-merge-commit merge: !4610 merge master_glm5 into master feat:Add adaptation support for GLM5.2 model Created-by: LinShua Commit-by: LinShua Merged-by: ascend-robot Description: ## What this PR does / why we need it? Add adaptation support for GLM5.2 model: 1.Revise and extend the model configuration mapping, add the indexer_types field, and optimize parameter mapping for MoE shared experts. 2.Add dedicated weight conversion scripts for GLM5.2 to support bidirectional checkpoint conversion between HF and Mcore formats, with preconfigured parallelism, MoE and MLA parameters. 3.Implement the share indexer capability, which can be enabled via the arguments --index-topk-freq and --index-skip-topk-offset. 4.Integrate data preprocessing into training scripts and provide supporting environment variable scripts. No breaking changes to external APIs. Internal tests have been completed, with normal inference and precision aligned with vllm-ascend. 5.The MLA function requires enabling the parameters --apply-rope-no-in-complex and --no-use-sparse-c8-indexer. ## Does this PR introduce any user-facing change? NA. ## How was this patch tested? 见PR See merge request: Ascend/MindSpeed-LLM!4610 | 3 个月前 | |
feat:Add adaptation support for GLM5.2 model Co-authored-by: LinShua<707894133@qq.com> # message auto-generated for no-merge-commit merge: !4610 merge master_glm5 into master feat:Add adaptation support for GLM5.2 model Created-by: LinShua Commit-by: LinShua Merged-by: ascend-robot Description: ## What this PR does / why we need it? Add adaptation support for GLM5.2 model: 1.Revise and extend the model configuration mapping, add the indexer_types field, and optimize parameter mapping for MoE shared experts. 2.Add dedicated weight conversion scripts for GLM5.2 to support bidirectional checkpoint conversion between HF and Mcore formats, with preconfigured parallelism, MoE and MLA parameters. 3.Implement the share indexer capability, which can be enabled via the arguments --index-topk-freq and --index-skip-topk-offset. 4.Integrate data preprocessing into training scripts and provide supporting environment variable scripts. No breaking changes to external APIs. Internal tests have been completed, with normal inference and precision aligned with vllm-ascend. 5.The MLA function requires enabling the parameters --apply-rope-no-in-complex and --no-use-sparse-c8-indexer. ## Does this PR introduce any user-facing change? NA. ## How was this patch tested? 见PR See merge request: Ascend/MindSpeed-LLM!4610 | 3 个月前 | |
feat(mcore): update GLM-5.2 A3 scripts Co-authored-by: LinShua<707894133@qq.com> # message auto-generated for no-merge-commit merge: !5000 merge glm52_sh_fix into master feat(mcore): update GLM-5.2 A3 scripts Created-by: LinShua Commit-by: LinShua Merged-by: ascend-robot Description: Fixes #1817 ## What this PR does / why we need it? 1. Align the GLM-5.2 744B A3 generation example by removing the default fused lightning indexer, sparse flash attention, and MLA absorb flags that are outside this script's intended configuration. 2. Simplify the GLM-5.2 744B A3 pretraining example and POC script by removing unused argument definitions and no longer passing the extra ROPE, memory, and profiling groups by default where applicable. 3. Add a GLM-5.2 210B 4K A3 pretraining POC script with reusable distributed, model, MLA, MoE, DSA, recomputation, and pipeline argument groups. This PR changes shell entrypoints only and does not modify training runtime code or public APIs. ## Does this PR introduce any user-facing change? Yes. Users gain a GLM-5.2 210B 4K A3 pretraining POC entrypoint under tests/poc/glm52/, while the existing 744B A3 example and POC commands no longer enable or pass the removed optional argument groups by default. ## How was this patch tested? 1. Ran changed-file pre-commit hooks on all four modified or added shell scripts. trailing-whitespace, end-of-file-fixer, check-added-large-files, check-merge-conflict, detect-private-key, codespell, and typos passed; file-type-specific YAML, JSON, Python, and C/C++ hooks were skipped. The repository-local gitleaks-offline-scan hook could not run because the configured ./gitleaks executable is absent. 2. Ran D:\Git\Git\bin\bash.exe -n on all four scripts; it passed. 3. Ran git diff --check origin/master...HEAD; it passed. 4. Verified git rev-list --count origin/master..HEAD returns 1 after squashing. 5. The GitCode push reported remote Git Hooks Checking [PASSED]. 6. No model training, accuracy, performance, or end-to-end functional test was run. See merge request: Ascend/MindSpeed-LLM!5000 | 1 个月前 | |
feat(mcore): add GLM-5.2 A5 POC scripts Co-authored-by: LinShua<707894133@qq.com> # message auto-generated for no-merge-commit merge: !5013 merge master_glm52_A5_sh into master feat(mcore): add GLM-5.2 A5 POC scripts Created-by: LinShua Commit-by: LinShua Merged-by: ascend-robot Description: Fixes #1825 ## What this PR does / why we need it? 1. Replaces the existing GLM-5.2 32B A5 POC entry with a 51B 4K FP8 configuration using TP1/PP2/EP4 parallelism, Transformer Engine MXFP8, index-topk=1024, fixed routing, the pipeline layout, and a matching log name. 2. Adds a GLM-5.2 210B 4K A5 FP8 POC entry with an eight-node TP1/PP2/EP32 configuration, Transformer Engine MXFP8 options, and portable data/tokenizer/checkpoint path placeholders. 3. Aligns --lr-warmup-iters to 500 in the GLM-5.2 744B A3 example, 744B A3 POC, and 210B A3 POC scripts. 4. This change is limited to five shell launch configurations and does not modify training runtime code or public APIs. ## Does this PR introduce any user-facing change? Yes. Users can run the new tests/poc/glm52/pretrain_glm52_210b_4k_A5_fp8_ptd.sh entry and the renamed tests/poc/glm52/pretrain_glm52_51b_4k_A5_ptd.sh entry. The 51B entry now uses the A5 MXFP8 and TP1/PP2/EP4 configuration, and the related A3 scripts now default to 500 learning-rate warmup iterations. ## How was this patch tested? 1. bash -n passed for all five changed shell scripts. 2. Changed-file pre-commit passed trailing-whitespace, end-of-file, added-large-file, merge-conflict, private-key, codespell, and typos checks. YAML, JSON, Python, and C/C++ hooks had no matching files. The repository-local gitleaks-offline-scan hook could not run because the required ./gitleaks executable is not present in the repository or local environment. 3. git diff --cached --check passed before amend. 4. The remote Git hook passed during git push --force-with-lease. 5. Functional, accuracy, performance, and end-to-end NPU training tests were not run. See merge request: Ascend/MindSpeed-LLM!5013 | 30 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 3 个月前 | ||
| 3 个月前 | ||
| 1 个月前 | ||
| 30 天前 |