Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.
InplaceApplyAdagradDA 在 Ascend950 FP16 场景下存在舍入边界精度误差。测试侧报告 22 个用例的 output0(var_out)MARE 超限;本地对已捕获输入的 L1_010、L1_128、L1_570 三例进行逐元素分析后确认:
问题由三层舍入误差叠加产生:
CAST_ROUND
CAST_RINT
以 L1_010 最差点为例,正确 FP32 商为 -5.3942203521728516e-06,恰好位于 FP16 0x805a 和 0x805b 的 midpoint;约 1 个 FP32 ULP 的上游误差即可改变最终 FP16 舍入结果。
-5.3942203521728516e-06
0x805a
0x805b
ttk_venv
c210daaf
use_locking
PYTHONHASHSEED=0
seed=42
stat_rel_err
修复前可观察到 output0 在 midpoint 邻域出现相邻 1 FP16 ULP 差异,近零 lane 的 MARE 超过阈值;output1、output2 不受影响。
修复前定位结果:
0.0109078
0.0159546
0.0201297
/assign
Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.
Describe the current behavior / 问题描述 (Mandatory / 必填)
InplaceApplyAdagradDA 在 Ascend950 FP16 场景下存在舍入边界精度误差。测试侧报告 22 个用例的 output0(var_out)MARE 超限;本地对已捕获输入的 L1_010、L1_128、L1_570 三例进行逐元素分析后确认:
问题由三层舍入误差叠加产生:
CAST_ROUND,midpoint tie 结果与 golden 的 round-to-nearest-even 不一致;CAST_RINT后,默认 FP32 Sqrt 的数 ULP 误差仍可能将结果推过 FP16 舍入边界;以 L1_010 最差点为例,正确 FP32 商为
-5.3942203521728516e-06,恰好位于 FP160x805a和0x805b的 midpoint;约 1 个 FP32 ULP 的上游误差即可改变最终 FP16 舍入结果。Environment / 环境信息 (Mandatory / 必填)
ttk_venv)c210daafSteps to reproduce the issue / 重现步骤 (Mandatory / 必填)
use_locking和标量约束构造 FP16 输入;固定PYTHONHASHSEED=0、seed=42、单设备单进程运行。stat_rel_err比较三个输出。修复前可观察到 output0 在 midpoint 邻域出现相邻 1 FP16 ULP 差异,近零 lane 的 MARE 超过阈值;output1、output2 不受影响。
Describe the expected behavior / 预期结果 (Mandatory / 必填)
Related log / screenshot / 日志 / 截图 (Mandatory / 必填)
修复前定位结果:
0.0109078、0.0159546、0.0201297。Special notes for this issue/备注 (Optional / 选填)