| ci: add permissions and PR labeling capability to pipeline Co-authored-by: BOCHENGZHANG<zhangbocheng6@h-partners.com> # message auto-generated for no-merge-commit merge: !178 merge ci-pr-labels-upstream into master ci: add permissions and PR labeling capability to pipeline Created-by: BOCHENGZHANG Commit-by: BOCHENGZHANG Merged-by: ascend-robot Description: ## 变更说明 - 权限扩展:新增 project: read、pr: write、note: write、action: read(保留原有 repository: read、id-token: write),用于打标签能力 - stage1 新增 associate_pr_comment 作业,通过 cann/.gitcode/actions/pre-pipeline 在流水线启动时给 PR 打标签(apply_label: 'true') - 新增 post 段,通过 cann/.gitcode/actions/query-pipeline 在流水线结束时根据结果给 PR 打标签 - 打标签动作认证使用 secrets.GIT_TOKEN;org_name 环境变量设为 Ascend,与 Ascend/op-plugin 的 PR-pipeline 实现对齐 - 流水线触发方式保持不变(pr_comment / pull_request_comment 评论 /compile 触发) ## 验证情况 已在 ComputingActionTest/MindSpeed-Ops 完成全流程验证: - 流水线变更合入测试仓(PR !5) - 通过 PR !6 评论 /compile 触发完整流水线,SCA、Antipoison、pre-commit、UT×4 全部通过 - pre_pipeline 启动打标签、query_pipeline 结束打标签均正常 ## 影响范围 仅修改 .gitcode/workflows/pr-pipeline-mindspeed-ops.yml,不涉及业务代码。 Related to [#81](https://gitcode.com/Ascend/MindSpeed-Ops/issues/81) See merge request: Ascend/MindSpeed-Ops!178 | 7 天前 |
| feat(atb): add MatmulAdd operators for fused wgrad accumulation Co-authored-by: Liz<lizhi166@huawei.com> # message auto-generated for no-merge-commit merge: !182 merge master into master feat(atb): add MatmulAdd operators for fused wgrad accumulation Created-by: Liz_ Commit-by: Liz Merged-by: ascend-robot Description: ## What this PR does / why we need it? Megatron gradient accumulation updates the main weight-gradient buffer as grad += grad_output.T @ total_input. The previous decomposed path materializes an [N, K] intermediate and launches a separate add kernel, increasing allocation and memory traffic. This PR adapts the MindSpeed MatmulAdd implementation to MindSpeed-Ops and provides: - npu_matmul_add_fp32 for bf16/fp16 operands accumulated into an fp32 grad; it uses the ATB fused MatmulAdd path on A2/A3 and addmm_ on A5. - npu_matmul_add_fp16 for bf16/fp16 operands and a matching half-precision grad; it replaces matmul + add_ with one in-place addmm_ call. - zero-sized-input protection, chip dispatch, builder integration, public documentation, and expanded unit coverage. - reproducible ATK case definitions and execution plugin for accuracy, device performance, and device memory validation. Generated cases, reports, and logs stay ignored; the final result summary is retained at tests/atk_tests/atb/matmul_add/atk_output/console/SUMMARY.md. This is feature enablement, not a bug fix. Related issue: N/A. ## Does this PR introduce any user-facing change? Yes. NPU callers can import npu_matmul_add_fp32 or npu_matmul_add_fp16 from mindspeed_ops.api.atb.npu_matmul_add. Both APIs mutate grad in place and return None. Inputs must be 2-D tensors with shapes total_input=[M,K], grad_output=[M,N], and grad=[N,K]. Operand dtypes must match and be bf16 or fp16. The fp32 API requires an fp32 grad buffer; the fp16 API requires grad to match the operand dtype. Empty inputs are a no-op. Usage and benchmark documentation: docs/zh/atb/matmul_add.md. ATK reproduction guide: tests/atk_tests/atb/matmul_add/README.md. ## How was this patch tested? Unit test in the 018 environment: - conda activate lz_py312_018 - python -m pytest tests/unit_tests/atb/test_npu_matmul_add.py -q -rs - Result: 8 passed, 7 skipped. The skipped cases require cann-nnal/ATB and were skipped because ATB_HOME_PATH was not set in this environment. ATK 26.4.30 was run on Ascend 910B3 (A2) with CANN 9.0.0 + ATB, Python 3.10, torch 2.7.1, and torch_npu 2.7.1.post4. Each API used 8 cases covering {bf16, fp16} operands across four shape tuples. - Accuracy, NPU vs CPU high-precision reference: fp32 8/8 Pass; fp16 8/8 Pass. - Device performance, fused dev0 vs decomposed dev1: fp32 average ratio 1.3662 with 100% case pass rate; fp16 average ratio 1.0007 with 87.5% case pass rate and overall Pass. - Device memory, fused dev0 vs decomposed dev1: fp32 and fp16 both 8/8 Pass with 100% memory pass rate under the fused/reference <= 1.1 criterion. The A2/A3 fp32 path requires cann-nnal and ATB_HOME_PATH. Its first invocation JIT-builds the C++ extension and takes about 30 seconds; performance measurements must use warmed, cache-hit runs. The public APIs are Ascend NPU-specific and do not provide a CUDA fallback. [RFC](https://gitcode.com/Ascend/MindSpeed-Ops/issues/80) See merge request: Ascend/MindSpeed-Ops!182 | 7 天前 |
| feat: support catlass Co-authored-by: zhuweichen<calvin_zhu0210@outlook.com> # message auto-generated for no-merge-commit merge: !131 merge feat/catlass-build-support into master feat: support catlass Created-by: Pr0Wh1teGive Commit-by: zhuweichen Merged-by: ascend-robot Description: ## What this PR does / why we need it? This PR adds common build support for CATLASS operators in MindSpeed-Ops: Adds the opt-in switch MINDSPEED_BUILD_CATLASS=1. Automatically detects and normalizes the target SoC. Automatically downloads a pinned CATLASS revision, while supporting local source through CATLASS_SOURCE_DIR. Adds an isolated Bisheng build process and packages the generated shared libraries. Adds a private BF16 BasicMatmul operator to verify the CATLASS build and runtime path. This is a feature PR rather than a bugfix. ## Does this PR introduce any user-facing change? CATLASS compilation can now be enabled with: export MINDSPEED_BUILD_CATLASS=1 pip install -e . --no-build-isolation --no-deps The default build behavior is unchanged when the option is disabled. Documentation: README.md. ## How was this patch tested? Tested on Atlas A3 with CANN, PyTorch, and torch_npu: CATLASS kernel compilation and shared-library loading succeeded. BF16 128 × 128 BasicMatmul was compared with torch.matmul. The observed maximum absolute error was 0. The regular build without CATLASS enabled was also verified. Limitations: the smoke operator only supports contiguous BF16 2D inputs, with M/N/K aligned to 16. See merge request: Ascend/MindSpeed-Ops!131 | 1 个月前 |
| docs: 移除Markdown文档中中英文及中文与数字之间的空格 Co-authored-by: EliasKaslan<eliaskaslan@outlook.com> # message auto-generated for no-merge-commit merge: !176 merge docs/remove-spaces-cn-en into master docs: 移除Markdown文档中中英文及中文与数字之间的空格 Created-by: EliasKaslan Commit-by: EliasKaslan Merged-by: ascend-robot Description: ## What this PR does / why we need it? 统一MindSpeed Ops文档空格规范:移除所有Markdown文档中中文与英文之间、中文与数字之间的空格。 - 覆盖范围:README、docs/zh、docker、ci、tests、tools等目录共35个Markdown文档,合计2346处空格调整 - 仅删除中英文、中文与数字之间的空格;代码块、行内代码、链接URL、表格结构均保持不变 - 相关章节标题同步调整后,已确认仓库内引用的锚点链接无失效(install_guide.md中被引用的章节标题原本即无空格) - docs/zh/release_notes_ops.md的空格调整由"版本说明按模板重构"的关联PR统一覆盖,避免两个PR改动冲突 ## Does this PR introduce any user-facing change? 无,仅文档格式统一,不涉及任何功能变更。 ## How was this patch tested? CI verification only(文档门禁markdownlint与链接有效性检查)。 See merge request: Ascend/MindSpeed-Ops!176 | 7 天前 |
| feat(triton): add layer_norm_fwd_1pass operator for Ascend NPU | 6 天前 |
| feat(triton): add layer_norm_fwd_1pass operator for Ascend NPU | 6 天前 |
| feat: 新增gitleaks敏感信息检测 Co-authored-by: wujinyuan1<wujinyuan1@huawei.com> # message auto-generated for no-merge-commit merge: !81 merge master into master feat: 新增gitleaks敏感信息检测 Created-by: wujinyuan1 Commit-by: wujinyuan1 Merged-by: ascend-robot Description: ## What this PR does / why we need it? 1. 引入gitleaks二进制离线扫描工具 2. 新增pre-commit/.gitleaks.toml配置,继承官方全部检测规则 3. 配置pre-commit钩子,提交前自动扫描密钥硬编码风险。 ## Does this PR introduce any user-facing change? 无. ## How was this patch tested? PR流水线pre-commit检测新增敏感信息检测. See merge request: Ascend/MindSpeed-Ops!81 | 3 个月前 |
| feat(triton): add layer_norm_fwd_1pass operator for Ascend NPU | 6 天前 |
| docs: 移除Markdown文档中中英文及中文与数字之间的空格 Co-authored-by: EliasKaslan<eliaskaslan@outlook.com> # message auto-generated for no-merge-commit merge: !176 merge docs/remove-spaces-cn-en into master docs: 移除Markdown文档中中英文及中文与数字之间的空格 Created-by: EliasKaslan Commit-by: EliasKaslan Merged-by: ascend-robot Description: ## What this PR does / why we need it? 统一MindSpeed Ops文档空格规范:移除所有Markdown文档中中文与英文之间、中文与数字之间的空格。 - 覆盖范围:README、docs/zh、docker、ci、tests、tools等目录共35个Markdown文档,合计2346处空格调整 - 仅删除中英文、中文与数字之间的空格;代码块、行内代码、链接URL、表格结构均保持不变 - 相关章节标题同步调整后,已确认仓库内引用的锚点链接无失效(install_guide.md中被引用的章节标题原本即无空格) - docs/zh/release_notes_ops.md的空格调整由"版本说明按模板重构"的关联PR统一覆盖,避免两个PR改动冲突 ## Does this PR introduce any user-facing change? 无,仅文档格式统一,不涉及任何功能变更。 ## How was this patch tested? CI verification only(文档门禁markdownlint与链接有效性检查)。 See merge request: Ascend/MindSpeed-Ops!176 | 7 天前 |
| 【feat】修改完善pre-commit配置文件 Co-authored-by: wujinyuan1<wujinyuan1@huawei.com> | 4 个月前 |
| feat(atb): add MatmulAdd operators for fused wgrad accumulation Co-authored-by: Liz<lizhi166@huawei.com> # message auto-generated for no-merge-commit merge: !182 merge master into master feat(atb): add MatmulAdd operators for fused wgrad accumulation Created-by: Liz_ Commit-by: Liz Merged-by: ascend-robot Description: ## What this PR does / why we need it? Megatron gradient accumulation updates the main weight-gradient buffer as grad += grad_output.T @ total_input. The previous decomposed path materializes an [N, K] intermediate and launches a separate add kernel, increasing allocation and memory traffic. This PR adapts the MindSpeed MatmulAdd implementation to MindSpeed-Ops and provides: - npu_matmul_add_fp32 for bf16/fp16 operands accumulated into an fp32 grad; it uses the ATB fused MatmulAdd path on A2/A3 and addmm_ on A5. - npu_matmul_add_fp16 for bf16/fp16 operands and a matching half-precision grad; it replaces matmul + add_ with one in-place addmm_ call. - zero-sized-input protection, chip dispatch, builder integration, public documentation, and expanded unit coverage. - reproducible ATK case definitions and execution plugin for accuracy, device performance, and device memory validation. Generated cases, reports, and logs stay ignored; the final result summary is retained at tests/atk_tests/atb/matmul_add/atk_output/console/SUMMARY.md. This is feature enablement, not a bug fix. Related issue: N/A. ## Does this PR introduce any user-facing change? Yes. NPU callers can import npu_matmul_add_fp32 or npu_matmul_add_fp16 from mindspeed_ops.api.atb.npu_matmul_add. Both APIs mutate grad in place and return None. Inputs must be 2-D tensors with shapes total_input=[M,K], grad_output=[M,N], and grad=[N,K]. Operand dtypes must match and be bf16 or fp16. The fp32 API requires an fp32 grad buffer; the fp16 API requires grad to match the operand dtype. Empty inputs are a no-op. Usage and benchmark documentation: docs/zh/atb/matmul_add.md. ATK reproduction guide: tests/atk_tests/atb/matmul_add/README.md. ## How was this patch tested? Unit test in the 018 environment: - conda activate lz_py312_018 - python -m pytest tests/unit_tests/atb/test_npu_matmul_add.py -q -rs - Result: 8 passed, 7 skipped. The skipped cases require cann-nnal/ATB and were skipped because ATB_HOME_PATH was not set in this environment. ATK 26.4.30 was run on Ascend 910B3 (A2) with CANN 9.0.0 + ATB, Python 3.10, torch 2.7.1, and torch_npu 2.7.1.post4. Each API used 8 cases covering {bf16, fp16} operands across four shape tuples. - Accuracy, NPU vs CPU high-precision reference: fp32 8/8 Pass; fp16 8/8 Pass. - Device performance, fused dev0 vs decomposed dev1: fp32 average ratio 1.3662 with 100% case pass rate; fp16 average ratio 1.0007 with 87.5% case pass rate and overall Pass. - Device memory, fused dev0 vs decomposed dev1: fp32 and fp16 both 8/8 Pass with 100% memory pass rate under the fused/reference <= 1.1 criterion. The A2/A3 fp32 path requires cann-nnal and ATB_HOME_PATH. Its first invocation JIT-builds the C++ extension and takes about 30 seconds; performance measurements must use warmed, cache-hit runs. The public APIs are Ascend NPU-specific and do not provide a CUDA fallback. [RFC](https://gitcode.com/Ascend/MindSpeed-Ops/issues/80) See merge request: Ascend/MindSpeed-Ops!182 | 7 天前 |
| feature: 引入 markdownlint 做文档语法检查,清理存量问题 Co-authored-by: wangjiangben<wangjiangben@huawei.com> # message auto-generated for no-merge-commit merge: !150 merge master into master feature: 引入 markdownlint 做文档语法检查,清理存量问题 Created-by: wangjiangben Commit-by: wangjiangben Merged-by: ascend-robot Description: # feature: 引入 markdownlint 做文档语法检查,清理存量问题 ## 概述 为仓库引入 **markdownlint-cli**(pre-commit 集成)做 Markdown 语法检查,并一次性清理存量问题,npx markdownlint-cli --config .markdownlint.yaml "**/*.md" 现为 0 issues。 裁剪了与中文正文、原生 HTML、多阶段标题冲突的风格规则,只保留真语法/结构检查;配置文件 .markdownlint.yaml 可被 VS Code markdownlint 扩展直接读取,本地编辑即时提示。 ## 变更清单 ### 新增 | 文件 | 说明 | |------|------| | .markdownlint.yaml | lint 配置:默认规则 + 规则裁剪 | | .markdownlintignore | 忽略清单(node_modules、.npm-cache、docs/zh/_menu.md) | ### 修改 | 文件 | 说明 | |------|------| | .pre-commit-config.yaml | 新增 markdownlint hook(rev v0.47.0) | | README.md | 失效锚点 #版本配套表 → #版本说明;补表格尾竖线 | | docs/zh/install_guide.md | 非描述性链接 [Link] → [TorchNPU插件版本] | | docs/zh/introduction.md | 图片补 alt | | ci/CI_GUID.md 及 docs/zh/**、tests/unit_tests/README.md、tools/skills/** 共 12 个 md | --fix 空行/列表标记/裸链接清理 | | tools/skills/operator-performance-profile/SKILL.md | 修复 130+ 处 emoji 乱码、14 处内联代码块、front matter 未转义引号 | ## 规则裁剪 | 关闭的规则 | 原因 | |------|------| | MD013 行宽 | 中文 + 宽表格 | | MD025 | 手册型文档多 H1 | | MD033 | <term>/<table> 原生 HTML | | MD036 | **表 1** 加粗题注 | | MD040 | shell/text 代码块未标注语言 | | MD060 | 表格 compact 风格不强制对齐 | | MD024(放宽 siblings_only) | 允许不同章节同名标题 | 保留语法关键规则:MD011/MD034/MD042/MD051(链接)、MD022/MD031/MD032(空行包裹)、MD009/MD047(尾随空格/末尾换行)、MD012/MD007/MD029/MD030(空行/列表)等。 ## 验证 bash npx markdownlint-cli --config .markdownlint.yaml "**/*.md" # 结果:0 issues 基线从约 1600 条违规收敛到 0:关闭风格类规则 → 约 230 条结构类问题,--fix 清约 230 处,手工修 4 处语义问题,重写 SKILL.md 乱码,最终归零。 ## 相关 - RFC:[[RFC]: MindSpeed-Ops 引入 markdownlint 做文档语法检查](https://gitcode.com/Ascend/MindSpeed-Ops/issues/64) See merge request: Ascend/MindSpeed-Ops!150 | 23 天前 |
| feature: 引入 markdownlint 做文档语法检查,清理存量问题 Co-authored-by: wangjiangben<wangjiangben@huawei.com> # message auto-generated for no-merge-commit merge: !150 merge master into master feature: 引入 markdownlint 做文档语法检查,清理存量问题 Created-by: wangjiangben Commit-by: wangjiangben Merged-by: ascend-robot Description: # feature: 引入 markdownlint 做文档语法检查,清理存量问题 ## 概述 为仓库引入 **markdownlint-cli**(pre-commit 集成)做 Markdown 语法检查,并一次性清理存量问题,npx markdownlint-cli --config .markdownlint.yaml "**/*.md" 现为 0 issues。 裁剪了与中文正文、原生 HTML、多阶段标题冲突的风格规则,只保留真语法/结构检查;配置文件 .markdownlint.yaml 可被 VS Code markdownlint 扩展直接读取,本地编辑即时提示。 ## 变更清单 ### 新增 | 文件 | 说明 | |------|------| | .markdownlint.yaml | lint 配置:默认规则 + 规则裁剪 | | .markdownlintignore | 忽略清单(node_modules、.npm-cache、docs/zh/_menu.md) | ### 修改 | 文件 | 说明 | |------|------| | .pre-commit-config.yaml | 新增 markdownlint hook(rev v0.47.0) | | README.md | 失效锚点 #版本配套表 → #版本说明;补表格尾竖线 | | docs/zh/install_guide.md | 非描述性链接 [Link] → [TorchNPU插件版本] | | docs/zh/introduction.md | 图片补 alt | | ci/CI_GUID.md 及 docs/zh/**、tests/unit_tests/README.md、tools/skills/** 共 12 个 md | --fix 空行/列表标记/裸链接清理 | | tools/skills/operator-performance-profile/SKILL.md | 修复 130+ 处 emoji 乱码、14 处内联代码块、front matter 未转义引号 | ## 规则裁剪 | 关闭的规则 | 原因 | |------|------| | MD013 行宽 | 中文 + 宽表格 | | MD025 | 手册型文档多 H1 | | MD033 | <term>/<table> 原生 HTML | | MD036 | **表 1** 加粗题注 | | MD040 | shell/text 代码块未标注语言 | | MD060 | 表格 compact 风格不强制对齐 | | MD024(放宽 siblings_only) | 允许不同章节同名标题 | 保留语法关键规则:MD011/MD034/MD042/MD051(链接)、MD022/MD031/MD032(空行包裹)、MD009/MD047(尾随空格/末尾换行)、MD012/MD007/MD029/MD030(空行/列表)等。 ## 验证 bash npx markdownlint-cli --config .markdownlint.yaml "**/*.md" # 结果:0 issues 基线从约 1600 条违规收敛到 0:关闭风格类规则 → 约 230 条结构类问题,--fix 清约 230 处,手工修 4 处语义问题,重写 SKILL.md 乱码,最终归零。 ## 相关 - RFC:[[RFC]: MindSpeed-Ops 引入 markdownlint 做文档语法检查](https://gitcode.com/Ascend/MindSpeed-Ops/issues/64) See merge request: Ascend/MindSpeed-Ops!150 | 23 天前 |
| ci: add GitCode PR pipeline Co-authored-by: BOCHENGZHANG<zhangbocheng6@h-partners.com> # message auto-generated for no-merge-commit merge: !154 merge master into master ci: add GitCode PR pipeline Created-by: BOCHENGZHANG Commit-by: BOCHENGZHANG Merged-by: ascend-robot Description: ## What this PR does / why we need it? Add a repository-specific GitCode PR pipeline based on the current MindSpeed-Ops CodeArts configuration. - Add CodeCheck and changed-file detection. - Add four sharded MindSpeed-Ops UT jobs using the confirmed Runner labels, container image, and existing CI entry script. - Support PR open/update/reopen events and the exact compile comment trigger for master and 26.1.0. - Add the Bandit TOML dependency required by the existing pre-commit configuration. ## Does this PR introduce any user-facing change? No. This change only adds GitCode CI configuration. ## How was this patch tested? The workflow was deployed to ComputingActionTest/MindSpeed-Ops and exercised through a validation PR. Trigger routing, Runner selection, container startup, merge-ref checkout, and UT entry-script execution were verified. The YAML diff and whitespace checks also passed locally. See merge request: Ascend/MindSpeed-Ops!154 | 21 天前 |
| feat: Add add_rms_norm_bias compiletion Co-authored-by: Liccol<740821011@qq.com> # message auto-generated for no-merge-commit merge: !54 merge aclnn-case into master feat: Add add_rms_norm_bias compiletion Created-by: Liccol Commit-by: Liccol Merged-by: ascend-robot Description: ## What this PR does / why we need it? 新增 add_rms_norm_bias 自定义 ACLNN 算子,将逐元素加法(Add)与 RMS 归一化合并为单个 AscendC Kernel 执行,用于 Transformer 模型推理中残差连接 + RMSNorm 的融合优化。 主要变更: - 新增 add_rms_norm_bias 算子源码 - 新增 cann/ 下的 ACLNN 绑定及 torch_binding 算子注册 - 新增 Python 调用接口(api/aclnn/)及 JIT 构建器(op_builder/) - 引入 cmake/common/scripts/utils 等编译工具链及公共组件 - 新增算子使用文档及开发指南(docs/aclnn/) ## Does this PR introduce any user-facing change? 是,新增 npu_add_rms_norm_bias 接口 ## How was this patch tested? 通过单元测试与 CPU 参考实现对比验证: - 对比分离执行 Add + RMSNorm 与融合算子的数值精度 - 覆盖 float32 / float16 / bfloat16 三种数据类型 - 覆盖有 beta / 无 beta 两种场景  See merge request: Ascend/MindSpeed-Ops!54 | 3 个月前 |
| fix: 商发wheel安全编译扫描整改(rpath/RELRO/栈保护/strip) Co-authored-by: Raining__<wangruining5@huawei.com> # message auto-generated for no-merge-commit merge: !173 merge master into master fix: 商发wheel安全编译扫描整改(rpath/RELRO/栈保护/strip) Created-by: Raining__ Commit-by: Raining__ Merged-by: ascend-robot Description: ## 背景 商发 wheel 包华为安全编译扫描整改。本 PR 修复扫描命中的 4 类问题,覆盖 wheel 内全部 host 侧 ELF 产物(pybind 扩展 + CATLASS 子库 + 设备 kernel .o)。设备 kernel .o 的 PIE/RELRO/栈保护/FS/strip 项因 CANN 设备二进制格式所限不适用,单独走豁免说明。 ## 修改明细 ### 1. 禁止 rpath(c781b17,另含 037ba4c+a57dfd5 的 kernel 侧修复) - mindspeed_ops/csrc/CMakeLists.txt:OPS_COMPILE_OPTIONS 追加 -fno-rtlib-add-rpath,消除 ccec 驱动默认 -frtlib-add-rpath 给设备 kernel .o 写入的 DT_RPATH - CMakeLists.txt:删除 mindspeed_ops_C 的 BUILD_RPATH/INSTALL_RPATH "$ORIGIN:$ORIGIN/lib" - mindspeed_ops/csrc/catlass/{basic_matmul,chunk_loss}/CMakeLists.txt:删除 INSTALL_RPATH "$ORIGIN" - mindspeed_ops/__init__.py:新增 _preload_catlass_libraries()(置于引导序列首位)——以绝对路径 ctypes.CDLL(..., RTLD_GLOBAL) 预载 mindspeed_ops/lib/*.so,利用 SONAME 注册机制使扩展的 NEEDED 依赖在不带 RPATH 的情况下正常解析,功能与原 $ORIGIN/lib RPATH 等价;与文件内既有 _preload_custom_opapi_library 同一模式,CATLASS 未使能(无 lib/ 目录)时为空操作 ### 2. GOT 重定位只读 (全 RELRO)与立即绑定(BIND_NOW)(39354c4、0706d44) - CMakeLists.txt:mindspeed_ops_C 追加链接选项 -Wl,-z,relro,-z,now - 两个 CATLASS 子库 target_link_options 追加 -Wl,-z,relro,-z,now(bisheng 默认仅 -z relro 部分 RELRO,补 -z now 到全 RELRO) - 验证:产物均出现 GNU_RELRO 段 + DT_BIND_NOW + FLAGS_1 NOW ### 3. 栈保护(3de5382) - CMakeLists.txt:mindspeed_ops_C 追加编译选项 -fstack-protector-all(选 -all 而非 -strong,确保所有函数插桩、确定性产生 __stack_chk_fail 引用) - 验证:产物动态符号表出现 __stack_chk_fail@GLIBC_2.4 ### 4. strip(fe649b5) - 两个 CATLASS 子库 target_link_options 追加 -Wl,-s - 验证:.symtab/.strtab 去除,.dynsym 完整保留(动态链接不受影响),file 判定 stripped ## 验证 - 每轮均在 py3.12 + CANN 9.1.0 标准构建镜像内全量重建 wheel,解包对全部 10 个 ELF 逐个 readelf 核验,已修复项无回归(rpath 全为 0) - 容器内 pip install 后 import mindspeed_ops 正常,CATLASS 预载链路日志确认生效 ## 兼容性 - 全部为编译/链接选项与装载顺序调整,无 API/行为变更;-fstack-protector-all 有常规量级运行时开销(每函数 canary 检查),对 host 侧绑定层可忽略 - 不带 CATLASS 的构建(MINDSPEED_BUILD_CATLASS=0)不受影响 See merge request: Ascend/MindSpeed-Ops!173 | 10 天前 |
| docs: add contribution guide (CONTRIBUTING.md & CONTRIBUTING_en.md) Co-authored-by: EliasKaslan<eliaskaslan@outlook.com> # message auto-generated for no-merge-commit merge: !161 merge docs/add-contributing-guide into master docs: add contribution guide (CONTRIBUTING.md & CONTRIBUTING_en.md) Created-by: EliasKaslan Commit-by: EliasKaslan Merged-by: ascend-robot Description: ## What this PR does / why we need it? Add contribution guide (CONTRIBUTING.md & CONTRIBUTING_en.md). ## Does this PR introduce any user-facing change? No. See merge request: Ascend/MindSpeed-Ops!161 | 7 天前 |
| docs: add contribution guide (CONTRIBUTING.md & CONTRIBUTING_en.md) Co-authored-by: EliasKaslan<eliaskaslan@outlook.com> # message auto-generated for no-merge-commit merge: !161 merge docs/add-contributing-guide into master docs: add contribution guide (CONTRIBUTING.md & CONTRIBUTING_en.md) Created-by: EliasKaslan Commit-by: EliasKaslan Merged-by: ascend-robot Description: ## What this PR does / why we need it? Add contribution guide (CONTRIBUTING.md & CONTRIBUTING_en.md). ## Does this PR introduce any user-facing change? No. See merge request: Ascend/MindSpeed-Ops!161 | 7 天前 |
| add ci& add ut Co-authored-by: shiyuan680<917935075@qq.com> | 5 个月前 |
| docs(atb): add custom operator development guide Co-authored-by: Liz<lizhi166@huawei.com> # message auto-generated for no-merge-commit merge: !185 merge master into master docs(atb): add custom operator development guide Created-by: Liz_ Commit-by: Liz Merged-by: ascend-robot Description: ## What this PR does / why we need it? add custom operator development guide https://gitcode.com/Ascend/MindSpeed-Ops/issues/80 ## Does this PR introduce any user-facing change? NAN ## How was this patch tested? CI PASSED See merge request: Ascend/MindSpeed-Ops!185 | 6 天前 |
| Fix docker build files and move torch to 2.10/cann to 9.1.0/triton-ascend to 3.2.2 Co-authored-by: Liccol<lichenhan2@huawei.com> # message auto-generated for no-merge-commit merge: !113 merge add-docker into master Fix docker build files and move torch to 2.10/cann to 9.1.0/triton-ascend to 3.2.2 Created-by: Liccol Commit-by: Liccol Merged-by: ascend-robot Description: ## What this PR does / why we need it? 本 PR 将 docker/ 目录下镜像构建的软件版本配套整体升级,并修复镜像构建流程。 **注意**:CI目录下的配套暂不进行修改,等CI机器上长跑验证后再另起PR合入,本PR仅关注镜像构建部分。 ### 软件版本配套 | 组件 | 旧版本 | 新版本 | |------|--------|--------| | CANN | 9.0.0-beta.2 | **9.1.0** | | PyTorch | 2.7.1 | **2.10.0** | | torch-npu | 2.7.1 | **2.10.0** | | triton-ascend | 3.2.1 | **3.2.2** | | Python | 3.11 | **3.12** | ### docker/Dockerfile 优化 - 移除无实际作用的 DNS 配置块 - 移除构建成功后的错误 Usage 提示输出 - rm /tmp/configure_repo.sh 改为 rm -f /tmp/configure_repo.sh,避免文件不存在时构建中断 - Ubuntu 分支 apt-get install 增加 --no-install-recommends,减小镜像体积 - 补充系统依赖与CI流程拉齐:findutils、jq、numactl-devel、tar、vim、which - 安装 MindSpeed-Ops 时增加 --no-build-isolation,复用容器内已装好的 torch-npu / triton-ascend,避免构建隔离环境重复拉取依赖 ### 文档同步 - docker/OVERVIEW.md:版本配套表、参数默认值表、示例命令全部更新;容器名 mindspeed → mindspeed-ops ## Does this PR introduce any user-facing change? 不涉及。 **构建参数接口无变化**。image_build.sh 和 build.sh 的命令行参数保持不变。 默认镜像 Tag 命名规则不变:mindspeed-ops:{branch}-{npu_type}-{os}-py{python_version}-{arch} **默认值变更**:基础镜像版本(9.0.0-beta.2 → 9.1.0)、Python(3.11 → 3.12)、PyTorch(2.7.1 → 2.10.0)、torch-npu(2.7.1 → 2.10.0)、triton-ascend(3.2.1 → 3.2.2)。 ## How was this patch tested? ### 构建验证 分别在A2/A3节点上完成镜像构建测试,通过: A2:  A3:   ### 容器环境验证 bash # 创建容器 docker run -itd \ --name mindspeed-ops \ --privileged \ --network host \ --ipc=host \ -v /usr/local/Ascend/driver:/usr/local/Ascend/driver \ -v /usr/local/dcmi:/usr/local/dcmi \ -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \ -v /etc/ascend_install.info:/etc/ascend_install.info \ -v /home:/home \ -v /data:/data \ -v /mnt:/mnt \ mindspeed-ops:master-a3-openeuler24.03-py3.11-x86_64 bash # 这里切换成具体生成的镜像名称 # 进入容器 docker exec mindspeed-ops bash # 执行import测试 source /usr/local/Ascend/ascend-toolkit/set_env.sh python -c "import mindspeed_ops; print(\"mindspeed_ops: OK\")"  ### 单元测试 bash pytest tests/unit_tests A2:  **TODO注意**:tests/unit_tests/triton/test_chunk_gated_delta_rule.py::TestChunkGatedDeltaRule::test_chunk_gated_delta_rule_forward_backward[B1-T1121-H8-K128-V128-chunk_size64-use_qk_l2norm_in_kernelTrue-cu_seqlens[0, 112, 209, 240, 281, 489, 523, 566, 689, 721, 785, 837, 985, 1071, 1121]] 异常,需在A2环境复测后提单跟踪 A3:   arm下tests/unit_tests/triton/test_mhc_pre_only.py::TestMhcPreOnlyOperator::test_mhc_pre_only偶发异常,已知问题,重跑后正常 x86下正常  See merge request: Ascend/MindSpeed-Ops!113 | 1 个月前 |
| Fix docker build files and move torch to 2.10/cann to 9.1.0/triton-ascend to 3.2.2 Co-authored-by: Liccol<lichenhan2@huawei.com> # message auto-generated for no-merge-commit merge: !113 merge add-docker into master Fix docker build files and move torch to 2.10/cann to 9.1.0/triton-ascend to 3.2.2 Created-by: Liccol Commit-by: Liccol Merged-by: ascend-robot Description: ## What this PR does / why we need it? 本 PR 将 docker/ 目录下镜像构建的软件版本配套整体升级,并修复镜像构建流程。 **注意**:CI目录下的配套暂不进行修改,等CI机器上长跑验证后再另起PR合入,本PR仅关注镜像构建部分。 ### 软件版本配套 | 组件 | 旧版本 | 新版本 | |------|--------|--------| | CANN | 9.0.0-beta.2 | **9.1.0** | | PyTorch | 2.7.1 | **2.10.0** | | torch-npu | 2.7.1 | **2.10.0** | | triton-ascend | 3.2.1 | **3.2.2** | | Python | 3.11 | **3.12** | ### docker/Dockerfile 优化 - 移除无实际作用的 DNS 配置块 - 移除构建成功后的错误 Usage 提示输出 - rm /tmp/configure_repo.sh 改为 rm -f /tmp/configure_repo.sh,避免文件不存在时构建中断 - Ubuntu 分支 apt-get install 增加 --no-install-recommends,减小镜像体积 - 补充系统依赖与CI流程拉齐:findutils、jq、numactl-devel、tar、vim、which - 安装 MindSpeed-Ops 时增加 --no-build-isolation,复用容器内已装好的 torch-npu / triton-ascend,避免构建隔离环境重复拉取依赖 ### 文档同步 - docker/OVERVIEW.md:版本配套表、参数默认值表、示例命令全部更新;容器名 mindspeed → mindspeed-ops ## Does this PR introduce any user-facing change? 不涉及。 **构建参数接口无变化**。image_build.sh 和 build.sh 的命令行参数保持不变。 默认镜像 Tag 命名规则不变:mindspeed-ops:{branch}-{npu_type}-{os}-py{python_version}-{arch} **默认值变更**:基础镜像版本(9.0.0-beta.2 → 9.1.0)、Python(3.11 → 3.12)、PyTorch(2.7.1 → 2.10.0)、torch-npu(2.7.1 → 2.10.0)、triton-ascend(3.2.1 → 3.2.2)。 ## How was this patch tested? ### 构建验证 分别在A2/A3节点上完成镜像构建测试,通过: A2:  A3:   ### 容器环境验证 bash # 创建容器 docker run -itd \ --name mindspeed-ops \ --privileged \ --network host \ --ipc=host \ -v /usr/local/Ascend/driver:/usr/local/Ascend/driver \ -v /usr/local/dcmi:/usr/local/dcmi \ -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \ -v /etc/ascend_install.info:/etc/ascend_install.info \ -v /home:/home \ -v /data:/data \ -v /mnt:/mnt \ mindspeed-ops:master-a3-openeuler24.03-py3.11-x86_64 bash # 这里切换成具体生成的镜像名称 # 进入容器 docker exec mindspeed-ops bash # 执行import测试 source /usr/local/Ascend/ascend-toolkit/set_env.sh python -c "import mindspeed_ops; print(\"mindspeed_ops: OK\")"  ### 单元测试 bash pytest tests/unit_tests A2:  **TODO注意**:tests/unit_tests/triton/test_chunk_gated_delta_rule.py::TestChunkGatedDeltaRule::test_chunk_gated_delta_rule_forward_backward[B1-T1121-H8-K128-V128-chunk_size64-use_qk_l2norm_in_kernelTrue-cu_seqlens[0, 112, 209, 240, 281, 489, 523, 566, 689, 721, 785, 837, 985, 1071, 1121]] 异常,需在A2环境复测后提单跟踪 A3:   arm下tests/unit_tests/triton/test_mhc_pre_only.py::TestMhcPreOnlyOperator::test_mhc_pre_only偶发异常,已知问题,重跑后正常 x86下正常  See merge request: Ascend/MindSpeed-Ops!113 | 1 个月前 |
| feat(atb): add MatmulAdd operators for fused wgrad accumulation Co-authored-by: Liz<lizhi166@huawei.com> # message auto-generated for no-merge-commit merge: !182 merge master into master feat(atb): add MatmulAdd operators for fused wgrad accumulation Created-by: Liz_ Commit-by: Liz Merged-by: ascend-robot Description: ## What this PR does / why we need it? Megatron gradient accumulation updates the main weight-gradient buffer as grad += grad_output.T @ total_input. The previous decomposed path materializes an [N, K] intermediate and launches a separate add kernel, increasing allocation and memory traffic. This PR adapts the MindSpeed MatmulAdd implementation to MindSpeed-Ops and provides: - npu_matmul_add_fp32 for bf16/fp16 operands accumulated into an fp32 grad; it uses the ATB fused MatmulAdd path on A2/A3 and addmm_ on A5. - npu_matmul_add_fp16 for bf16/fp16 operands and a matching half-precision grad; it replaces matmul + add_ with one in-place addmm_ call. - zero-sized-input protection, chip dispatch, builder integration, public documentation, and expanded unit coverage. - reproducible ATK case definitions and execution plugin for accuracy, device performance, and device memory validation. Generated cases, reports, and logs stay ignored; the final result summary is retained at tests/atk_tests/atb/matmul_add/atk_output/console/SUMMARY.md. This is feature enablement, not a bug fix. Related issue: N/A. ## Does this PR introduce any user-facing change? Yes. NPU callers can import npu_matmul_add_fp32 or npu_matmul_add_fp16 from mindspeed_ops.api.atb.npu_matmul_add. Both APIs mutate grad in place and return None. Inputs must be 2-D tensors with shapes total_input=[M,K], grad_output=[M,N], and grad=[N,K]. Operand dtypes must match and be bf16 or fp16. The fp32 API requires an fp32 grad buffer; the fp16 API requires grad to match the operand dtype. Empty inputs are a no-op. Usage and benchmark documentation: docs/zh/atb/matmul_add.md. ATK reproduction guide: tests/atk_tests/atb/matmul_add/README.md. ## How was this patch tested? Unit test in the 018 environment: - conda activate lz_py312_018 - python -m pytest tests/unit_tests/atb/test_npu_matmul_add.py -q -rs - Result: 8 passed, 7 skipped. The skipped cases require cann-nnal/ATB and were skipped because ATB_HOME_PATH was not set in this environment. ATK 26.4.30 was run on Ascend 910B3 (A2) with CANN 9.0.0 + ATB, Python 3.10, torch 2.7.1, and torch_npu 2.7.1.post4. Each API used 8 cases covering {bf16, fp16} operands across four shape tuples. - Accuracy, NPU vs CPU high-precision reference: fp32 8/8 Pass; fp16 8/8 Pass. - Device performance, fused dev0 vs decomposed dev1: fp32 average ratio 1.3662 with 100% case pass rate; fp16 average ratio 1.0007 with 87.5% case pass rate and overall Pass. - Device memory, fused dev0 vs decomposed dev1: fp32 and fp16 both 8/8 Pass with 100% memory pass rate under the fused/reference <= 1.1 criterion. The A2/A3 fp32 path requires cann-nnal and ATB_HOME_PATH. Its first invocation JIT-builds the C++ extension and takes about 30 seconds; performance measurements must use warmed, cache-hit runs. The public APIs are Ascend NPU-specific and do not provide a CUDA fallback. [RFC](https://gitcode.com/Ascend/MindSpeed-Ops/issues/80) See merge request: Ascend/MindSpeed-Ops!182 | 7 天前 |