ffn/ffn_worker_batching/tests/ ├── assets/ │ ├── spec.py # TestSpec 适配器(精度容差 + 自定义比较逻辑) │ └── impl/ │ ├── inputs.py # 输入数据生成(customize_inputs) │ └── golden.py # CPU 参考实现(golden) └── ut/op_api/ ├── test_aclnn_ffn_worker_batching.csv # 基础用例集(25 个) └── test_aclnn_ffn_worker_batching_full.csv # 泛化用例集(200 个)
文件路径:tests/assets/spec.py
tests/assets/spec.py
作用:将 golden、customize_inputs、精度容差和自定义比较逻辑统一注册到一个 TestSpec 类中,供 TTK 框架在 ACLNN 模式下加载使用。
关键设计:
golden_module = _load_impl("golden") inputs_module = _load_impl("inputs") golden_module._DATA_CACHE = inputs_module._DATA_CACHE
inputs.py 在执行 customize_inputs 时会将生成的 CPU 测试数据缓存到 _DATA_CACHE 字典(key=testcase_name)。golden.py 的 golden 函数需要这些 CPU 数据来计算参考结果(因为 schedule_context tensor 中存的是 device 偏移量,CPU 侧无法解析)。通过 golden_module._DATA_CACHE = inputs_module._DATA_CACHE 将两个模块的缓存打通。
inputs.py
customize_inputs
_DATA_CACHE
golden.py
golden_module._DATA_CACHE = inputs_module._DATA_CACHE
tolerance = { "float16": {"standard": "binary_equal"}, # FP16: 二进制精确匹配 "bfloat16": {"standard": "stat_rel_err"}, # BF16: 统计相对误差(位模式可能有差异) "int8": {"standard": "binary_equal"}, # INT8: 二进制精确匹配 "float32": {"standard": "binary_equal"}, # float32: 二进制精确匹配 "int32": {"standard": "binary_equal"}, # int32: 二进制精确匹配 "int64": {"standard": "binary_equal"}, # int64: 二进制精确匹配 }
该算子是整数索引排序 + 原样字节搬运,无浮点运算,因此除 BF16 外均使用二进制精确匹配。BF16 使用统计相对误差是因为搬运过程中可能存在位模式差异。
kernel 只写入前 actual_token_num 个有效 token 的输出,其余位置保留 TTK 的初始化值(1)。golden 用 0 填充未使用位置。因此需要自定义比较逻辑,只比较有效区域:
actual_token_num
[0, 0]
BF16 tensor 的特殊处理:to_np 辅助函数将 BF16 tensor 先转 float32 再转 numpy,避免 np.asarray 不支持 BF16 的 TypeError。
to_np
np.asarray
__spec__ = {"aclnnFfnWorkerBatching": "AclnnFfnWorkerBatchingTestSpec"}
key 必须与 CSV 中的 api_name 完全一致。
api_name
文件路径:tests/assets/impl/inputs.py
tests/assets/impl/inputs.py
作用:为每个测试用例生成确定性输入数据,构建 schedule_context 结构体,将附属数据追加到 schedule_context storage 后面(相对偏移模式),并缓存 CPU 数据供 golden 使用。
def _seed_from_name(name): h = 0 for ch in name: h = (h * 131 + ord(ch)) & 0x7FFFFFFF return h
从用例名生成种子,确保同一用例每次运行数据一致。数据包括:
expert_ids
token_data
token_scales
session_ids
micro_batch_ids
masked
_MASK_VALUE
single_expert
问题:TTK 框架在 customize_inputs 阶段调用 aclrtMalloc 分配的 HBM,在后续 create_context 后对 kernel 不可见(返回 107000 错误)。
aclrtMalloc
create_context
方案:把 token_data / expert_ids / session_ids / micro_batch_ids 等附属数据追加到 schedule_context tensor 的 storage 后面(保持 view shape (1024,) 不变,只扩展 untyped_storage)。TTK 的 copy_torch_tensor_to_hbm 用 storage().nbytes() 拷贝完整 storage 到 device。
(1024,)
copy_torch_tensor_to_hbm
storage().nbytes()
schedule_context 中存相对偏移量(从 schedule_context 起始到数据起始的字节偏移),offset 640 写 TEST_MAGIC = 0x54455354 标志位。kernel 读取此标志后,将 buffer 指针字段解释为 schedule_context 地址 + 偏移。
TEST_MAGIC = 0x54455354
schedule_context 地址 + 偏移
_build_token_data_bytes 对 tokenDtype=1(BF16)先把 float32 转 torch.bfloat16 再取字节,确保 HBM 中是 BF16 位模式而非 float32。
_build_token_data_bytes
对 tokenDtype=2,把每个 token 的 int8 数据(H 字节)和 float32 scale(4 字节)连续排布为 [A, M, BS, K, H+4] 的 uint8 数组。
[A, M, BS, K, H+4]
_build_token_info_buf 生成 [A, M, F] 的 int32 数组,F = 2 + BSK。每个 [a, m] 的前两个元素是 flag=1(数据就绪)和 layer_id=0,后面是 BSK 个 expert_id。
_build_token_info_buf
[A, M, F]
[a, m]
用 scheduleContext.set_(new_storage, offset, shape, stride) 替换底层 storage,然后用 untyped_storage()[:total_size].copy_(...) 拷贝数据。
scheduleContext.set_(new_storage, offset, shape, stride)
untyped_storage()[:total_size].copy_(...)
文件路径:tests/assets/impl/golden.py
tests/assets/impl/golden.py
作用:模拟 kernel 的排序 + gather + group_list 逻辑,作为精度比对的参考实现。
通过 _get_cached_data(testcase_name) 从 _DATA_CACHE 获取 inputs.py 缓存的 CPU 数据。因为 schedule_context 中的偏移量在 CPU 侧无意义。
_get_cached_data(testcase_name)
_sort_expert_ids
EXPERT_MASK_VALUE
int32.max
np.argsort(kind='stable')
_generate_group_list
[expert_id, token_count]
按排序顺序遍历每个有效 token:
sorted_order[i]
gidx
a_idx = gidx // (BS*K)
bs_idx = (gidx % (BS*K)) // K
k_idx = gidx % K
session_ids_buf[a_idx]
micro_batch_ids_buf[a_idx]
session = a_idx
micro_batch = cur_micro_batch_id
token_data[a_idx, mb_idx, bs_idx, k_idx, :]
token_scales[a_idx, mb_idx, bs_idx, k_idx]
用 np.ones 而非 np.zeros。因为 tokenDtype≠2 时 kernel 不写 dynamic_scale 输出,TTK 把纯输出 tensor 初始化为 1,golden 需要匹配这个初始值。
np.ones
np.zeros
返回 8 个 torch tensor,顺序与 CSV 的 output_tensor_indexes 一致:y, group_list, session_ids, micro_batch_ids, token_ids, expert_offsets, dynamic_scale, actual_token_num。
output_tensor_indexes
文件路径:tests/ut/op_api/test_aclnn_ffn_worker_batching.csv
tests/ut/op_api/test_aclnn_ffn_worker_batching.csv
作用:25 个基础测试用例,覆盖算子的核心功能和主要参数组合。
testcase_name
aclnnFfnWorkerBatching_fp16_basic
__spec__
aclnnFfnWorkerBatching
tensor_view_shapes
((1024,),(72,128),(8,2),(72,),...)
tensor_dtypes
('int8','float16','int64','int32',...)
attributes
{'expertNum': 8, 'maxOutShape': [1,8,9,128], ...}
(1,2,3,4,5,6,7,8)
其中 Y = A × BS × K,E = expertNum。
文件路径:tests/ut/op_api/test_aclnn_ffn_worker_batching_full.csv
tests/ut/op_api/test_aclnn_ffn_worker_batching_full.csv
作用:200 个泛化测试用例,在基础用例集之上大幅扩展参数覆盖范围,用于全面验证算子在各 shape 组合下的正确性。
fp16_min_a1_bs1_k2_h1_e4_norm
fp16_min_a1_bs1_k2_h1_e4_recv
fp16_masked_h128
fp16_masked_h4096
bf16_masked_h256
int8_masked_h128
fp16_masked_a4_h512
fp16_single_expert_h128
fp16_single_expert_h4096
bf16_single_expert_h256
int8_single_expert_h128
fp16_layer0_h128
fp16_layer2_h128
fp16_layer0_h4096
fp16_a128_bs1_k9_h128_norm
fp16_a256_bs1_k2_h128_norm
fp16_a1_bs8_k2_h4096_norm
int8_k2_h512_norm
TTK 框架读取 CSV ↓ 对每行用例: 1. 加载 spec.py → AclnnFfnWorkerBatchingTestSpec 2. 调用 customize_inputs(inputs.py) → 修改 schedule_context tensor + 缓存 CPU 数据 3. TTK 将 schedule_context 拷贝到 HBM 4. 调用 aclnnFfnWorkerBatchingGetWorkspaceSize → 执行 tiling 5. 调用 aclnnFfnWorkerBatching → 执行 kernel 6. 调用 golden(golden.py) → 生成 CPU 参考结果 7. 调用 compare(spec.py) → 自定义精度比较 8. 输出 precision_status: PASS/FAIL
# 基础用例(25 个) cd /workspace/ops-test-kit python3 -m ttk aclnn \ -i .../test_aclnn_ffn_worker_batching.csv \ -o /tmp/result.csv \ --plugin .../spec.py \ --pc 1 --ti=0-25 # 泛化用例(200 个) python3 -m ttk aclnn \ -i .../test_aclnn_ffn_worker_batching_full.csv \ -o /tmp/result_full.csv \ --plugin .../spec.py \ --pc 1 --ti=0-199
python3 -c " import csv p=0;f=0 with open('/tmp/result.csv') as fh: for row in csv.DictReader(fh): if row.get('precision_status')=='PASS': p+=1 else: f+=1 print(f'Total:{p+f} PASS:{p} FAIL:{f}') "
/assign @cpy_123456
FfnWorkerBatching TTK 测试描述文档
1. 文件总览
2. spec.py — TestSpec 适配器
文件路径:
tests/assets/spec.py作用:将 golden、customize_inputs、精度容差和自定义比较逻辑统一注册到一个 TestSpec 类中,供 TTK 框架在 ACLNN 模式下加载使用。
关键设计:
2.1 模块加载与缓存共享
golden_module = _load_impl("golden") inputs_module = _load_impl("inputs") golden_module._DATA_CACHE = inputs_module._DATA_CACHEinputs.py在执行customize_inputs时会将生成的 CPU 测试数据缓存到_DATA_CACHE字典(key=testcase_name)。golden.py的 golden 函数需要这些 CPU 数据来计算参考结果(因为 schedule_context tensor 中存的是 device 偏移量,CPU 侧无法解析)。通过golden_module._DATA_CACHE = inputs_module._DATA_CACHE将两个模块的缓存打通。2.2 精度容差定义
tolerance = { "float16": {"standard": "binary_equal"}, # FP16: 二进制精确匹配 "bfloat16": {"standard": "stat_rel_err"}, # BF16: 统计相对误差(位模式可能有差异) "int8": {"standard": "binary_equal"}, # INT8: 二进制精确匹配 "float32": {"standard": "binary_equal"}, # float32: 二进制精确匹配 "int32": {"standard": "binary_equal"}, # int32: 二进制精确匹配 "int64": {"standard": "binary_equal"}, # int64: 二进制精确匹配 }该算子是整数索引排序 + 原样字节搬运,无浮点运算,因此除 BF16 外均使用二进制精确匹配。BF16 使用统计相对误差是因为搬运过程中可能存在位模式差异。
2.3 自定义 compare 函数
kernel 只写入前
actual_token_num个有效 token 的输出,其余位置保留 TTK 的初始化值(1)。golden 用 0 填充未使用位置。因此需要自定义比较逻辑,只比较有效区域:actual_token_num行(每行 H 个元素)[0, 0]终止符之前的行actual_token_num个元素actual_token_num个元素actual_token_num个元素actual_token_num个元素BF16 tensor 的特殊处理:
to_np辅助函数将 BF16 tensor 先转 float32 再转 numpy,避免np.asarray不支持 BF16 的 TypeError。2.4 注册
__spec__ = {"aclnnFfnWorkerBatching": "AclnnFfnWorkerBatchingTestSpec"}key 必须与 CSV 中的
api_name完全一致。3. inputs.py — 输入数据生成
文件路径:
tests/assets/impl/inputs.py作用:为每个测试用例生成确定性输入数据,构建 schedule_context 结构体,将附属数据追加到 schedule_context storage 后面(相对偏移模式),并缓存 CPU 数据供 golden 使用。
3.1 确定性随机数据生成
def _seed_from_name(name): h = 0 for ch in name: h = (h * 131 + ord(ch)) & 0x7FFFFFFF return h从用例名生成种子,确保同一用例每次运行数据一致。数据包括:
expert_ids:NORM 模式 shape=(A, BS, K),RECV 模式 shape=(A, M, BS, K)token_data:shape=(A, M, BS, K, H),dtype 取决于 tokenDtypetoken_scales:仅 INT8 量化时存在,shape=(A, M, BS, K),float32session_ids:shape=(A,),值为 0..A-1micro_batch_ids:shape=(A,),全 03.2 特殊用例数据变异
masked时,~15% 的 expert_id 被设为 1000000(_MASK_VALUE),验证 mask 过滤逻辑single_expert时,所有 expert_id 设为 0,验证退化场景3.3 数据布局与 HBM 可见性方案
问题:TTK 框架在
customize_inputs阶段调用aclrtMalloc分配的 HBM,在后续create_context后对 kernel 不可见(返回 107000 错误)。方案:把 token_data / expert_ids / session_ids / micro_batch_ids 等附属数据追加到 schedule_context tensor 的 storage 后面(保持 view shape
(1024,)不变,只扩展 untyped_storage)。TTK 的copy_torch_tensor_to_hbm用storage().nbytes()拷贝完整 storage 到 device。schedule_context 中存相对偏移量(从 schedule_context 起始到数据起始的字节偏移),offset 640 写
TEST_MAGIC = 0x54455354标志位。kernel 读取此标志后,将 buffer 指针字段解释为schedule_context 地址 + 偏移。3.4 BF16 数据转换
_build_token_data_bytes对 tokenDtype=1(BF16)先把 float32 转 torch.bfloat16 再取字节,确保 HBM 中是 BF16 位模式而非 float32。3.5 INT8 量化 scale 打包
对 tokenDtype=2,把每个 token 的 int8 数据(H 字节)和 float32 scale(4 字节)连续排布为
[A, M, BS, K, H+4]的 uint8 数组。3.6 token_info 构建(RECV 模式)
_build_token_info_buf生成[A, M, F]的 int32 数组,F = 2 + BSK。每个[a, m]的前两个元素是 flag=1(数据就绪)和 layer_id=0,后面是 BSK 个 expert_id。3.7 storage 扩展
用
scheduleContext.set_(new_storage, offset, shape, stride)替换底层 storage,然后用untyped_storage()[:total_size].copy_(...)拷贝数据。4. golden.py — CPU 参考实现
文件路径:
tests/assets/impl/golden.py作用:模拟 kernel 的排序 + gather + group_list 逻辑,作为精度比对的参考实现。
4.1 数据获取
通过
_get_cached_data(testcase_name)从_DATA_CACHE获取 inputs.py 缓存的 CPU 数据。因为 schedule_context 中的偏移量在 CPU 侧无意义。4.2 排序逻辑 (
_sort_expert_ids)EXPERT_MASK_VALUE(1000000) 的 expert_id 映射为int32.maxnp.argsort(kind='stable')),masked token 排到最后4.3 group_list 生成 (
_generate_group_list)[expert_id, token_count]行[0, 0]作为终止符[0, 0]4.4 gather 逻辑
按排序顺序遍历每个有效 token:
sorted_order[i]获取原始索引gidxa_idx = gidx // (BS*K),bs_idx = (gidx % (BS*K)) // K,k_idx = gidx % Ksession_ids_buf[a_idx]和micro_batch_ids_buf[a_idx]查表session = a_idx,micro_batch = cur_micro_batch_idtoken_data[a_idx, mb_idx, bs_idx, k_idx, :]提取 hidden statestoken_scales[a_idx, mb_idx, bs_idx, k_idx]4.5 dynamic_scale 初始化
用
np.ones而非np.zeros。因为 tokenDtype≠2 时 kernel 不写 dynamic_scale 输出,TTK 把纯输出 tensor 初始化为 1,golden 需要匹配这个初始值。4.6 输出
返回 8 个 torch tensor,顺序与 CSV 的
output_tensor_indexes一致:y, group_list, session_ids, micro_batch_ids, token_ids, expert_offsets, dynamic_scale, actual_token_num。5. test_aclnn_ffn_worker_batching.csv — 基础用例集
文件路径:
tests/ut/op_api/test_aclnn_ffn_worker_batching.csv作用:25 个基础测试用例,覆盖算子的核心功能和主要参数组合。
CSV 字段说明
testcase_nameaclnnFfnWorkerBatching_fp16_basicapi_name__spec__key 一致)aclnnFfnWorkerBatchingtensor_view_shapes((1024,),(72,128),(8,2),(72,),...)tensor_dtypes('int8','float16','int64','int32',...)attributes{'expertNum': 8, 'maxOutShape': [1,8,9,128], ...}output_tensor_indexes(1,2,3,4,5,6,7,8)9 个 tensor 的索引
其中 Y = A × BS × K,E = expertNum。
25 个用例分类
6. test_aclnn_ffn_worker_batching_full.csv — 泛化用例集
文件路径:
tests/ut/op_api/test_aclnn_ffn_worker_batching_full.csv作用:200 个泛化测试用例,在基础用例集之上大幅扩展参数覆盖范围,用于全面验证算子在各 shape 组合下的正确性。
参数覆盖范围
用例分布
用例组成(8 个部分)
边界用例说明
fp16_min_a1_bs1_k2_h1_e4_normfp16_min_a1_bs1_k2_h1_e4_recvfp16_masked_h128fp16_masked_h4096bf16_masked_h256int8_masked_h128fp16_masked_a4_h512fp16_single_expert_h128fp16_single_expert_h4096bf16_single_expert_h256int8_single_expert_h128fp16_layer0_h128fp16_layer2_h128fp16_layer0_h4096fp16_a128_bs1_k9_h128_normfp16_a256_bs1_k2_h128_normfp16_a1_bs8_k2_h4096_normint8_k2_h512_norm7. TTK 执行流程
执行命令
# 基础用例(25 个) cd /workspace/ops-test-kit python3 -m ttk aclnn \ -i .../test_aclnn_ffn_worker_batching.csv \ -o /tmp/result.csv \ --plugin .../spec.py \ --pc 1 --ti=0-25 # 泛化用例(200 个) python3 -m ttk aclnn \ -i .../test_aclnn_ffn_worker_batching_full.csv \ -o /tmp/result_full.csv \ --plugin .../spec.py \ --pc 1 --ti=0-199结果检查
python3 -c " import csv p=0;f=0 with open('/tmp/result.csv') as fh: for row in csv.DictReader(fh): if row.get('precision_status')=='PASS': p+=1 else: f+=1 print(f'Total:{p+f} PASS:{p} FAIL:{f}') "8. 测试验证结果