已关闭
A2/A3新增 GenericBlockSparseAttentionGrad 和 generic_block_sparse_attention_grad_metadata算子 #5110
tramp-ll创建于  16 天前关闭于  16 天前
tramp-ll成员
16 天前 创建

GenericBlockSparseAttentionGrad是通用块稀疏注意力的反向计算算子。依据sparseBlockIdx/sparseBlockCount(稀疏块索引表)定义的索引,仅在被选中的KV块上计算和传播梯度,支持动态、可变长的分块稀疏模式。
新增 GenericBlockSparseAttentionGrad 和 generic_block_sparse_attention_grad_metadata算子

likedislike
游震成员
16 天前 评论:

/assign @tramp-ll

likedislike
CANN-robotCANN-robot成员
16 天前 将 tramp-ll 设为负责人
weihao18成员
16 天前 评论:

/assign @tramp-ll

likedislike
CANN-robotCANN-robot成员
16 天前 关闭了 issue
CANN-robotCANN-robot成员
16 天前 添加了label:resolved