已关闭
[Requirement|需求建议]: NsaCompressAttentionInfer算子性能优化 #1104
dingxu创建于 3月9日关闭于 3月30日
3月9日 将 xtqh 设为负责人
3月9日 将 monologue815 设为负责人
3月9日 移除了负责人 xtqh
3月9日 将 L_Euler 设为负责人
3月9日 将 xdnjust 设为负责人
3月9日 移除了负责人 L_Euler
dingxu
3月12日 评论:
3月12日 评论:
针对该需求的优化算子,已提PR,链接:https://gitcode.com/cann/ops-transformer/pull/2615


monologue815
3月27日 评论:
3月27日 评论:
针对NSA算子性能优化,接纳


3月30日 关闭了 issue
3月30日 添加了label:resolved
Thanks for sending an requirement! Please fill in the following template to help quickly solve your problem.
Backgroud(背景信息)
本次需求对于功能无变化,意在提升importance score计算性能。
importance score计算主要有以下两个步骤
ptslc[j] = m=0∑dl′−1n=0∑dl−1ptcmp[dl′j − m − n]ptslc′=h=1∑Hptslc,(h)
当前NsaCompressAttentionInfer算子中对于importance score使用vector进行计算,会存在使用transpose,mul,add等操作,性能差。
期望通过提前生成系数矩阵,将importance score计算移入cube来提升性能
既有方案分析
分为两层循环
关于系数的说明:

以l = 32, l' = 64, d = 16为例,l'/d=4,l/d=2,则ptslc [1]由以下索引的数值组成
Benefit / Necessity (价值/作用)
通过矩阵方式计算importance score,提升性能明显。




优化前
优化后
以模型典型shape(q head num=16,kv head num=1,select block size = 64, compress block size = 32, compress stride = 16,paged block size = 128)为例,测试几组场景,单算子性能提升明显。
Design(设计方案)
参数矩阵的生成
本版本代码实现,为了不修改接口,使用kernel内部生成参数矩阵的方式
l = 32, l' = 64, d = 16为例,参数矩阵为矩阵乘方案
通过矩阵乘优化importance score计算

计算流程

def ImportanceScoreCube: Init result workspace to 0 DataCopy W from GM to L1 LoadData W from L1 to L0B # 参数右矩阵常驻L0B for startRow in range(0, rows, rowSplit): endRow = startRow + rowSplit for l1StartCol in range(0, cols, l1ColSplit): l1EndCol = l1StartCol + l1StartCol DataCopy P[startRow: endRow, l1StartCol: l1EndCol] from GM to L1 for l0startCol in range(0, l1ColSplit, l0ColSplit): LoadData from L1 to L0A # 左矩阵的分块 execute PW matmul set atomic add true fixpipe L0C to GM set atomic add falseQ:为什么系数矩阵需要前缀列
P矩阵的不重复load,而importance score的计算在多块之间有重叠。以P矩阵为32列,参数矩阵为16行,即分两次计算为例:此时计算第一块(1-16列)时,实际上会依赖于第P矩阵的第17列的值
当前通过fixpipe时增加atomic add开关实现。
Q:W矩阵如此稀疏,会不会导致性能变差?
通过后面的流水可知,矩阵计算为搬移运bound,mmad占比小,不是瓶颈。
Q:kernel内部生成参数矩阵会不会性能差
不会,生成参数矩阵在vector,完全被QK掩盖