已合并
[feat] 新增triton_dense_to_jagged算子atk测试迁移(npu部分) #123
[feat] 新增triton_dense_to_jagged算子atk测试迁移(npu部分) #123
已合并
QianZH97创建于 25 天前
QianZH97
QianZH97成员
25 天前

合入来源

问题/功能描述

[feat] 新增triton_dense_to_jagged算子atk测试迁移

修改方案描述

triton_dense_to_jagged 算子源自 FBGEMM,将 dense tensor 转为 jagged tensor。含 NPU 端 ATK 精度与性能测试、限制参数生成脚本与自定义 API 封装,支持 add/mul 元素级融合。

是否涉及UT/ST

开发自检

likedislike
Pull Request已成功合入, 合并人@ascend-robot
(感谢 QianZH97 的贡献)
QianZH97QianZH97成员
25 天前 关联了里程碑:ops-rec 26.2.0
atomgit-bot
atomgit-bot
25 天前 评论:

变更摘要

本 PR 为源自 FBGEMM 的 triton_dense_to_jagged 算子(将规则 dense tensor 转换为不规则 jagged tensor,支持 add/mul 元素级融合)新增 GPU/NPU 双端 ATK 精度与性能测试迁移。改动包含双端 Triton kernel 实现(NPU 端带 @triton.autotune 自动调优)、继承 BaseApi 并通过 @register("triton_dense_to_jagged") 注册的自定义 API 封装 TritonDenseToJaggedApi、基于 GENERATOR_REGISTRY 注册的参数约束生成器 DenseToJaggedGenerator,以及描述测试输入参数的 YAML 配置和性能测试 JSON 文件。

主要改动

  • 新增 Triton kernel 实现:在 gpu_atk_test/triton_dense_to_jagged.py 与 npu_atk_test/triton_dense_to_jagged.py 中新增 triton_dense_to_jagged 及辅助融合算子 tensor_elementwise_add / tensor_elementwise_mul(通过 operation_function 常量表达式选择融合方式);NPU 版本额外增加 @triton.autotune 配置(16/32/64/128/256 的线程块组合),GPU 版本则固定 32×32 线程块。
  • 新增 ATK API 封装:双端新增 TritonDenseToJaggedApi,实现 init_by_input_data(获取输入参数、按 dense 实际 shape 确定性生成满足约束的 jagged_offsets、创建输出 tensor 并计算 kernel 参数)与 __call__(启动 kernel 并支持 NPU/CUDA 同步);内部辅助函数 _generate_jagged_offsets 处理 3D/4D dense 对应的 1 个或 2 个 offset,_jagged_offsets_to_dense_indice 计算 jagged 索引到 dense 索引的映射。
  • 新增参数约束生成器:npu_atk_test/generate_triton_dense_to_jagged.py 中的 DenseToJaggedGenerator.after_case_config 对用例参数做约束修正,包括特征维度 D <= 8192、3D/4D dense 的 shape 关系(B < seq < D、B1 < B2 < seq < D)、offset 从 0 开始且递增、相邻差值及末值边界限制,以及 operation_jagged_value 与 dense 数据类型一致。
  • 新增 ATK 用例配置:npu_atk_test/triton_dense_to_jagged.yaml 定义测试用例输入,包括 dense(fp32/fp16/bf16,3/4 维)、jagged_offsets(int32/int64,tuple 数量 1/2)、operation_function(null/"add"/"mul")与 operation_jagged_value 标量,并指定精度评估与性能评估配置。
  • 新增性能测试文件:npu_atk_test/triton_dense_to_jagged_performance.json 作为性能用例配置文件加入(当前为空文件)。
likedislike
不准确?
atomgit-bot
atomgit-bot
25 天前 评论:

代码审查

审查结论

已审查文件

  1. experimental/triton/atk_test/triton_dense_to_jagged/gpu_atk_test/triton_dense_to_jagged.py(Triton kernel)
    • 已逐行审查 kernel 主循环、mask、stride 计算与融合分支。happy path(连续 dense + 确定性生成的 offsets)经手工推演正确,报告 2 个问题(block_offset 增量、-1 融合语义)。
  2. experimental/triton/atk_test/triton_dense_to_jagged/gpu_atk_test/triton_dense_to_jagged_api.py(ATK API 封装)
    • 已审查 offset 确定性生成、dense_indices 计算、kernel 启动参数与 call 流程,报告 3 个问题(空 offsets 校验、裸 except、operation_function 大小写)。

发现统计

  • P0:0
  • P1:0
  • P2:1(kernel 内层循环 block_offset 增量未乘 dense_col_stride,非连续 dense 输入时列寻址错误)
  • P3:4(jagged_offsets 为空导致 IndexError、get_device 裸 except 吞异常、operation_function 大小写未归一化导致静默走 mul、dense_indice=-1 时 mul 融合将 jagged 数据清零)

总体风险评估

中低风险。该 PR 是从 FBGEMM 移植的 Triton kernel + ATK 测试封装,核心转换逻辑在测试用例的生成输入(连续 dense、约束内 offset)下经推演正确,未发现会导致 P0/P1 级别的 happy-path 运行错误或数据损坏问题。报告的问题集中在潜在触发条件(非连续 dense、缺失/非标准 yaml 参数、-1 哨兵路径)与健壮性/校验缺失,当前用例下均被掩盖,但属于可明确触发的逻辑缺陷或缺陷防御,建议修复后再合入。


审查结论

各变更文件审查情况

  1. generate_triton_dense_to_jagged.py — 已审查。发现与 yaml generate 字段的注册名不一致问题(见下,归属 yaml 文件锚点);文件内其余约束修正逻辑存在若干防御性缺口,但多被 API 侧的 offset 重生成覆盖,未单独上报。
  2. triton_dense_to_jagged.py — 已审查。kernel 为 FBGEMM 上游迁移代码,参数映射、block 掩码、offset 寻址逻辑经核对无新增缺陷。
  3. triton_dense_to_jagged.yaml — 已审查。发现 generate: 与生成器注册名不一致(P2);键名拼写无错误。
  4. triton_dense_to_jagged_api.py — 已审查。发现 3 个问题:_jagged_offsets_to_dense_indice 越界判断 off-by-one(P2)、4D dense 配 1 个 offset 未处理(P2)、融合时 operation_jagged_value 缺失导致 None 指针算术(P3)。
  5. triton_dense_to_jagged_performance.json — 已审查。发现严重问题:内容整体指向 triton_jagged_to_dense 算子(P1)。

按优先级统计

  • P1:1 项 — 性能测试配置文件错配为兄弟算子 triton_jagged_to_dense 的内容(复制粘贴错误),dense_to_jagged 的性能测试将无法运行或跑错算子。
  • P2:3 项 — ① _jagged_offsets_to_dense_indice 第 231 行 idx > dense_sizes[0] 应为 >=,导致所有 4D 用例的越界行获得非法正索引并越界读 dense 内存;② yaml generate 名称与生成器注册名(generate_dense_to_jagged)不一致,约束修正逻辑可能不生效;③ 4D dense 配 1 个 offset 的组合(yaml 允许,约 1/4 概率)未被处理,kernel 按 3D 模式运行导致结果错误。
  • P3:1 项 — operation_function 为 add/mul 而 operation_jagged_value 缺失时,kernel 对 None 指针做算术崩溃。

总体风险判断

本 PR 为 ATK 测试迁移,kernel 主体为上游可靠代码,但外围的 API 封装与配置文件存在多处真实缺陷:性能配置错配会使性能测试失效(P1),_jagged_offsets_to_dense_indice 的越界判断 bug 会使全部 4D 精度用例产生错误结果/越界内存访问(P2),生成器注册名不一致及 4D+1offset 组合未处理进一步影响用例有效性。建议在合入前修复 P1/P2 项(尤其性能 JSON 错配与越界 off-by-one),否则该迁移交付的测试无法正确、完整地运行。

类型 数量
🔴 阻塞 1
🟡 建议 0

⛔ 需要修改

likedislike
不准确?
ascend-robot
ascend-robot成员
25 天前 评论:
流水线 PR-pipeline_ops-rec#609 [ commitID:b7f8bb6d ] 已完成
>>>代码风格自动修复执行成功(无修复内容)
阶段 任务名 状态 详情
编译构建 PR-build-SAST-check ✅ COMPLETED >>>
Build-2.7.1_arm ✅ COMPLETED >>>
Build-2.10.0_arm ✅ COMPLETED >>>
Build-2.7.1_X86 ✅ COMPLETED >>>
Build-2.10.0_X86 ✅ COMPLETED >>>
恶意代码检查 Antipoison ✅ COMPLETED >>>
编码安全与规范检查 pre-commit ✅ COMPLETED >>>
clang-tidy ✅ COMPLETED >>>
开源片段检查 SCA ✅ COMPLETED >>>
开发者测试 presmoke ✅ COMPLETED >>>
流水线 PR-pipeline_ops-rec ✅ COMPLETED >>>
此流水线已支持下列评论快捷指令,仅PR创建者和白名单成员评论有效
  • compile : 运行流水线
  • retry : 重试流水线所有失败子任务
  • retry <任务名> : 仅重试指定失败子任务
  • stop : 停止流水线
likedislike
ascend-robotascend-robot成员
25 天前 添加了label:ascend-cla/yes
此处折叠了137条消息 查看更多
li-da-xiong成员
23 天前 评论:

/approve

likedislike
ascend-robotascend-robot成员
23 天前 添加了label:approved
ascend-robotascend-robot成员
23 天前 合入了pull request
ascend-robot
ascend-robot成员
23 天前 评论:

Pull Request 已合并或已关闭。

If you want to solve this problem, you can click here to do it in the FAQs.

likedislike
ascend-robot
ascend-robot成员
23 天前 评论:

Pull Request 已合并或已关闭。

If you want to solve this problem, you can click here to do it in the FAQs.

likedislike