Thanks for sending an requirement! Please fill in the following template to help quickly solve your problem.
优化四个场景的性能(新增2个模板,修改2个模板) 对于大batch小cp场景,使用前缀和+simt实现。 对于切repeat场景,使用输出分核策略+前缀和获取偏移。 解决小cp轴场景性能差和核间负载不均的问题。
NA
优化repeat_interleave算子部分场景性能
计算全局exclusive前缀和: 每个核内前缀和的第一个位置为哨兵0,搬出的时候偏移1,使得每个位置前缀和不包含自身(exclusive),同时记录核内总和放在workspace上coreSumGm_; SyncAll(); 搬入coreSumGm_,计算每个核前面核的总和,加到当前核的核内前缀和上;(第0核不用处理)
查找各核处理的repeats的开始和结尾: 使用simt开2个线程在GM上进行二分查找头和尾。
simt计算开3维线程: // 最外层循环对应当前处理的batch个数 for (batch = threadIdx.z; batch < curBatchCount; batch += blockDim.z) { urBatchXOffset = inputBatchOffset * batch; curBatchYOffset = outputBatchOffset * batch; // 第二层循环为repeat个数 for (repeatIdx = threadIdx.y; repeatIdx < totalRepeats; repeatIdx += blockDim.y) { curRepeatNum = repeatsGm[repeatIdx]; inputX = xGm[curBatchXOffset + repeatIdx * cpNum + threadIdx.x]; // 内存循环为当前repeat的复制次数 for (repeat = 0; repeat < curRepeatNum; repeat += 1) { curOutRepeatOffset = prefixSumGm[repeatIdx]; // 输出偏移直接从前缀和获取 yGm[curBatchYOffset + (curOutRepeatOffset + repeat) * cpNum + threadIdx.x] = inputX; } } }
/assign
Thanks for sending an requirement! Please fill in the following template to help quickly solve your problem.
Backgroud(背景信息)
优化四个场景的性能(新增2个模板,修改2个模板)
对于大batch小cp场景,使用前缀和+simt实现。
对于切repeat场景,使用输出分核策略+前缀和获取偏移。
解决小cp轴场景性能差和核间负载不均的问题。
Origin(信息来源)
NA
Benefit / Necessity (价值/作用)
优化repeat_interleave算子部分场景性能
Design(设计方案)
计算全局exclusive前缀和:
每个核内前缀和的第一个位置为哨兵0,搬出的时候偏移1,使得每个位置前缀和不包含自身(exclusive),同时记录核内总和放在workspace上coreSumGm_;
SyncAll();
搬入coreSumGm_,计算每个核前面核的总和,加到当前核的核内前缀和上;(第0核不用处理)
查找各核处理的repeats的开始和结尾:
使用simt开2个线程在GM上进行二分查找头和尾。
simt计算开3维线程:
// 最外层循环对应当前处理的batch个数
for (batch = threadIdx.z; batch < curBatchCount; batch += blockDim.z) {
urBatchXOffset = inputBatchOffset * batch;
curBatchYOffset = outputBatchOffset * batch;
// 第二层循环为repeat个数
for (repeatIdx = threadIdx.y; repeatIdx < totalRepeats; repeatIdx += blockDim.y) {
curRepeatNum = repeatsGm[repeatIdx];
inputX = xGm[curBatchXOffset + repeatIdx * cpNum + threadIdx.x];
// 内存循环为当前repeat的复制次数
for (repeat = 0; repeat < curRepeatNum; repeat += 1) {
curOutRepeatOffset = prefixSumGm[repeatIdx]; // 输出偏移直接从前缀和获取
yGm[curBatchYOffset + (curOutRepeatOffset + repeat) * cpNum + threadIdx.x] = inputX;
}
}
}