已关闭
[Feature] npu_weight_quant_preprocess 支持 A16W4(INT4/FP4/MXFP4)紧凑排布数据流 #4358
马琦钧创建于  8月25日关闭于  8月27日
马琦钧
马琦钧成员
8月25日 创建

背景

WeightQuantBatchMatmulV2 在 Ascend950 上新增 A16W4(INT4/FP4/MXFP4)4-bit 紧凑排布(uint8 载体)weight 支持(ops-nn 侧见 cann/ops-nn#8545),需要 PTA 侧 npu_weight_quant_preprocess 配套支持对应数据流。

当前 npu_weight_quant_preprocess 仅支持 MX A8W4 / int8 等存量流程,不支持 A16S4(INT4 per-tensor/per-channel/per-group)与 A16F4(FP4 per-group/MX)的 4-bit 紧凑 weight 预处理。

目标

  1. A16S4 INT4:per-tensor / per-channel / per-group 全场景支持。
  2. A16F4 FP4:per-group / MX 支持。
  3. torch 接口保持原 9 参形式,无新增参数。
  4. ND/NZ 内部分流:per-tensor 一律 ND 直拷(不支持 NZ);per-channel / per-group 非转置 → NZ_C0_16 转换,转置 → ND 直拷,均不支持则报错。
  5. FORMAT_FAKE_TO_REAL 映射扩展:NZ_C0_8→NZ_C0_16(A16 4-bit),保留 NZ_C0_16→NZ_C0_32(A8)。
  6. npu_weight_quant_batchmatmul 支持 uint8 载体 4-bit ND weight 按打包方向还原逻辑 K/N;is_weight_nz_4bit_compact 识别 NZ_C0_8/NZ_C0_16。

验收标准

  • eager:per-tensor/per-channel/per-group、转置自动 ND、负向用例通过,与 CPU golden bit-exact。
  • 图模式(torchair 静态+动态)全数据流用例通过。
  • 存量 MX A8W4 / int8 流程回归不受影响。

关联 PR

likedislike
马琦钧马琦钧成员
8月25日 关联了看板:FrameworkPTAdapter 版本issue看板
TorchNPU-BotTorchNPU-Bot成员
8月25日 添加了label:triage-review
TorchNPU-Bot
TorchNPU-Bot成员
8月25日 评论:

issue待分派,添加triage-review标签

likedislike
ascend-robotascend-robot成员
8月25日 添加了label:feature
马琦钧马琦钧成员
8月25日 关联了pull request:docs: 补充 npu_weight_quant_preprocess 与 npu_weight_quant_batchmatmul 的 A16W4 数据流文档
马琦钧
马琦钧成员
8月27日 评论:

/close

likedislike
ascend-robot
ascend-robot成员
8月27日 评论:

Notice

@maqijun , this issue is currently under CVE service control and cannot be closed directly. Please comment /check-issue.

likedislike
马琦钧
马琦钧成员
8月27日 评论:

/close

likedislike
ascend-robot
ascend-robot成员
8月27日 评论:

Notice

@maqijun , this issue is currently under CVE service control and cannot be closed directly. Please comment /check-issue.

likedislike
马琦钧马琦钧成员
8月27日 关闭了 issue