已关闭
[Feature] npu_weight_quant_preprocess 支持 A16W4(INT4/FP4/MXFP4)紧凑排布数据流 #4358
马琦钧创建于 8月25日关闭于 8月27日
8月25日 关联了看板:FrameworkPTAdapter 版本issue看板
8月25日 添加了label:triage-review
TorchNPU-Bot
8月25日 评论:
8月25日 评论:
issue待分派,添加triage-review标签


8月25日 添加了label:feature
8月25日 关联了pull request:docs: 补充 npu_weight_quant_preprocess 与 npu_weight_quant_batchmatmul 的 A16W4 数据流文档
8月25日 关联了pull request:docs: 补充 npu_weight_quant_preprocess 与 npu_weight_quant_batchmatmul 的 A16W4 数据流文档
马琦钧
8月27日 评论:
8月27日 评论:
/close


ascend-robot
8月27日 评论:
8月27日 评论:
Notice
@maqijun , this issue is currently under CVE service control and cannot be closed directly. Please comment /check-issue.


马琦钧
8月27日 评论:
8月27日 评论:
/close


ascend-robot
8月27日 评论:
8月27日 评论:
Notice
@maqijun , this issue is currently under CVE service control and cannot be closed directly. Please comment /check-issue.


8月27日 关闭了 issue
背景
WeightQuantBatchMatmulV2 在 Ascend950 上新增 A16W4(INT4/FP4/MXFP4)4-bit 紧凑排布(uint8 载体)weight 支持(ops-nn 侧见 cann/ops-nn#8545),需要 PTA 侧
npu_weight_quant_preprocess配套支持对应数据流。当前
npu_weight_quant_preprocess仅支持 MX A8W4 / int8 等存量流程,不支持 A16S4(INT4 per-tensor/per-channel/per-group)与 A16F4(FP4 per-group/MX)的 4-bit 紧凑 weight 预处理。目标
npu_weight_quant_batchmatmul支持 uint8 载体 4-bit ND weight 按打包方向还原逻辑 K/N;is_weight_nz_4bit_compact识别 NZ_C0_8/NZ_C0_16。验收标准
关联 PR