已关闭
[Bug]: 静态编译在原地修改算子场景存在精度问题 #800
rich创建于 11 天前关闭于 10 天前
11 天前 添加了label:bug
11 天前 添加了label:bug
11 天前 修改了issue 的描述
11 天前 关联了pull request:fix: avoid double execution in static kernel compilation
11 天前 将 rich9527 设为负责人
10 天前 关闭了 issue
10 天前 issue状态由 TODO 改变为 DONE
10 天前 添加了label:resolved
在提交问题之前,请通过搜索现有和历史问题确保该问题尚未被提出并解决。
您的环境信息
静态编译会单算子执行一遍图,以获取算子JSON,但在模型包含原地修改算子时会出现副作用(模型多执行一遍,输入会改变)
🐛 请描述bug
import torch import torch_npu import torch.nn as nn class InplaceModel(nn.Module): def forward(self, x): x.add_(1) return x * 2 def main(): if not hasattr(torch, "npu") or not torch.npu.is_available(): raise RuntimeError("This reproducer requires an Ascend NPU and torch_npu.") model = InplaceModel().npu() compiled = torch.compile( model, backend="npugraph_ex", fullgraph=True, dynamic=False, options={ "static_kernel_compile": True, # Force the dump path on every run so the repro is deterministic. "disable_static_kernel_compile_cache": True, }, ) original = torch.arange(4, dtype=torch.float32, device="npu") actual_input = original.clone() print("input before call:", actual_input.cpu()) expected = (original + 1) * 2 actual = compiled(actual_input) torch.npu.synchronize() print("expected:", expected.cpu()) print("actual: ", actual.cpu()) print("input after call:", actual_input.cpu()) torch.testing.assert_close(actual_input, original + 1) if not torch.allclose(actual, expected): raise AssertionError( "Regression: the fx graph was executed more than once in the static-kernel dump step." ) print("PASS") if __name__ == "__main__": main()