已合并
[sync] PR-39802: [fix]profiler fix config cache when analyse multi card #39873
[sync] PR-39802: [fix]profiler fix config cache when analyse multi card #39873
已合并
ascend-robot创建于 7月2日
ascend-robot
ascend-robot成员
7月2日

1. Origin pull request:

https://gitcode.com/Ascend/pytorch/merge_requests/39802

https://gitcode.com/Ascend/pytorch/issues/2573

Sha Datetime Message
74141fe8 2026-07-01 19:36:24 +0800 CST [fix]profiler fix config cache when analyse multi card
likedislike
Pull Request已成功合入, 合并人@ascend-robot
(感谢 ascend-robot 的贡献)
ascend-robotascend-robot成员
7月2日 创建了 pull request,commit 8ef55017
ascend-robotascend-robot成员
7月2日 关联了issue:[Bug]: 离线解析设置的max_process_number的值小于总ascend_pt数量时,解析出来db文件的rank与profiler_info.json的rank不一致
ascend-robotascend-robot成员
7月2日 添加了label:ascend-cla/yes
ascend-robot
ascend-robot成员
7月2日 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here
You can get sig-info at here


PR Approval Progress

Congratulations! All modules have met the lgtm and approve requirements.

Module Approval Details

module lgtm status approve status
test 陈豪, renyujin (2/2) 陈豪 (1/1)
torch_npu/profiler 陈豪, renyujin (2/2) 陈豪 (1/1)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

ascend-ds-bot, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
ascend-robotascend-robot成员
7月2日 添加了label:ci-pipeline-running
ascend-robot
ascend-robot成员
7月2日 评论:

ascend docs pipeline is running...

likedislike
ascend-robotascend-robot成员
7月2日 添加了label:docs-ci-pipeline-running
ascend-robot
ascend-robot成员
7月2日 评论:

✅ 跳过 docs ci 检查,没有需要检查的文档文件

likedislike
ascend-robotascend-robot成员
7月2日 删除了label:docs-ci-pipeline-running
ascend-robotascend-robot成员
7月2日 添加了label:docs-ci-pipeline-success
ascend-robotascend-robot成员
7月2日 删除了label:ci-pipeline-running
ascend-robotascend-robot成员
7月2日 添加了label:ci-pipeline-passed
ascend-robot
ascend-robot成员
7月2日 评论:
流水线 PR-pipeline_pytorch#40694 [ commitID:d7350137 ] 已完成
>>>代码风格自动修复执行成功(无修复内容)
阶段 任务名 状态 详情
编译构建 Build_X86 >>>
Build_ARM >>>
Build_LibTorch_x86 >>>
Build_LibTorch_ARM >>>
Build_X86_torchair 🛑 >>>
Build_ARM_torchair 🛑 >>>
patch_test 🛑 >>>
恶意代码检查 Antipoison >>>
编码安全与规范检查 CodeCheck >>>
check_error >>>
CodeCheck_lintrunner >>>
开源片段检查 SCA >>>
开发者测试 UT_X86_Part_01 🛑 >>>
UT_X86_Part_02 🛑 >>>
UT_ARM_A3_Part_01 🛑 >>>
UT_ARM_A3_Part_02 🛑 >>>
UT_ARM_A2_Part_01 >>>
UT_ARM_A2_Part_02 >>>
UT_ARM_A2_Part_03 >>>
UT_inductor_Part_01 🛑 >>>
UT_inductor_Part_02 🛑 >>>
UT_inductor_Part_03 🛑 >>>
UT_inductor_Part_04 🛑 >>>
UT_DIST_ARM_Part_01 🛑 >>>
UT_DIST_ARM_Part_02 🛑 >>>
UT_DIST_ARM_Part_03 🛑 >>>
UT_DIST_ARM_Part_04 🛑 >>>
UT_ARM_A2_Select_Part_01 >>>
UT_ARM_A2_Select_Part_02 >>>
流水线 PR-pipeline_pytorch >>>
此流水线已支持下列评论快捷指令,仅PR创建者和白名单成员[wujinyuan1, huangjingwei, liangsongwei, yashi999, culechan, Dring, wuyouqi1, L1919_snow, qq_52711437, WhiteNight12, nomiz, xiu_21, ffmh, wanglijun55, hss-shuai, husichao, smallsilly, lanshaozuishuai, jimmyisme1, lzy0920232, alpha-junh, Sunshine_Youngster, wei_zhuoyi, zhangyihuiben, zyw-hw, zzzkeke, rmch, yangch0324, LucciC, AACAES, renyujin, wjlflyer, senzhen-town, pengjingyou, qsc97, limuan, yule100, xiaoqi-zhou, kuhn7, chenxingying, hanye02, zichun_ye, anyrenwei, kkjocker, wangzili121, Lu_G, yvjc, puddingfjz, HandsoemLemon, bigprestigee1, huawuyi, zhenyu10, dairenjie, du-jin-hang, zou-jieyu, adelaideliu, TrHan, wanlinan, Windwindzzz, pengqihw, kisnwang, yuheng_wang, honghao_wang, jizewei, zhangguoguang, sunyu-xuan, chenrayray, hbhu_bin, liujunzhu, c_34, LiNuoh, maoyuanpeng1, zzhongmin, zhaoyu65, bellatan, jiabaolin, zhuofanshen, wencaiwen, lu_zhuge, caoshuyang, molly12, lyx324521, LQ1206, gitcode-chenjiao, cai-weiwei1989, CHDong, ogqin, yuanlipingGit, xuqinglin1, lqz2, zouwei1, chaoluoa, paradox325, jackzhang1116, yaoyao, akh, yujiacheng, dengjie0116, Hubert11111, Shine_Ws, wslhj555, longqiand, OYtao666, JiaqingQiang, luyyyy, Kingbelial, zhanghaiyu0101, wenxp1018, yanliu-luoluo, ksun_sekiro, liyong328, wgzheng, tangky, vivi_is_coding, aoiaoisola, weixin_44494597, wangmengmengwang65667, hid57809721, qq_35468730, comeonup, C547032, gcw_m5OQChA4, yao_yao_ling_xian, cnnbwcy, szqfes_12, cora_19, cann_lilin, can, shawnylee233]评论有效
  • compile、compile_inductor、compile_torchair : 运行流水线
  • retry : 重试流水线所有失败子任务
  • retry <任务名> : 仅重试指定失败子任务
  • stop : 停止流水线
likedislike
atomgit-bot
atomgit-bot
7月2日 评论:

变更摘要

此 PR 修复了 profiler 在多卡分析场景下 ProfilerConfig 配置缓存未正确重置的问题。在多卡分析时,ProfilerConfig_is_load 标志在上一次分析后仍为 True,导致后续卡的分析无法重新加载配置信息。修复方案为新增 reset_is_load() 方法,并在每次解析 profiling 数据前调用该方法以清除缓存标志,确保每张卡的配置信息能被正确加载。

主要改动

  • 新增 ProfilerConfig.reset_is_load() 方法:在 _profiler_config.py 中为 ProfilerConfig 类添加了 reset_is_load() 方法,将 _is_load 标志重置为 False,允许后续重新加载配置。
  • 解析前重置配置加载状态:在 _profiling_parser.pyProfilingParser 解析流程中,于 ProfilerConfig().load_info() 调用前插入 ProfilerConfig().reset_is_load() 调用,确保每张卡的配置信息都能被加载而不受上一张卡缓存的影响。
  • 新增单元测试:在 test_profiler_config.py 中添加了 test_reset_is_load_will_reset_is_load_to_false 测试用例,验证 reset_is_load() 方法能正确将 _is_load 置为 False 且不影响其他属性(如 _rank_id)。
likedislike
atomgit-bot
atomgit-bot
7月2日 评论:

代码审查

我已经仔细审查了所有三个变更文件的 diff 和上下文代码。以下是每个文件的分析:

  • torch_npu/profiler/analysis/_profiler_config.pyProfilerConfig@Singleton 装饰的类。新增的 reset_is_load() 方法将 _is_load 重置为 False,从而允许 load_info() 在后续调用时重新读取配置(而不是因 _is_load 守卫而直接返回)。

  • torch_npu/profiler/analysis/_profiling_parser.py — 在 analyse_profiling_data 中的 load_info() 调用之前添加了 ProfilerConfig().reset_is_load()。由于 ProfilerConfig 是单例,两个 ProfilerConfig() 调用返回的是同一个实例。此变更确保每个 profiling 路径的配置都会被重新读取,而不是使用缓存的配置。

  • test/profiler/analysis/test_profiler_config.py — 新增的测试方法 test_reset_is_load_will_reset_is_load_to_false 验证了 reset_is_load() 能正确重置 _is_load 而不影响 _rank_id

我检查了以下潜在问题,但均 未发现 此 diff 引入了真正的问题:

  1. 正确性reset_is_load() + load_info() 的组合对于每个 worker 进程中顺序处理多张卡的情况是正确且安全的。load_info 会覆盖所有配置字段(_rank_id_export_type_activities 等)。由于使用 concurrent.futures.ProcessPoolExecutor(多进程,非多线程),不同 worker 进程之间不存在共享内存,因此不存在竞态条件。

  2. 部分加载风险:如果 load_info 在中途失败(例如 _get_json_data 返回 {} 或某个子方法抛出异常),_is_load 保持为 False,下次调用时会重新尝试。字段回退到默认值或当前值,这是预先存在的设计行为,并非此 diff 引入。

  3. 字段残留load_syscnt_infoload_timediff_info 在 JSON 缺少某些字段时会回退到当前值(例如 self._freq)。这可能导致来自上一张卡的某些值被保留——但这是这些加载方法预先存在的回退行为,并非此 diff 引入。此 diff 实际上通过允许 load_info 重新运行,改善了原本完全不重新加载任何字段的问题。

  4. 测试隔离性:由于 ProfilerConfig 是单例,所有测试共享同一个实例——这是现有测试套件的一个预先存在的模式,并非此 diff 引入。

总结

  • P0–P2 问题:0
  • P3 问题:0
  • 总体结论:此 diff 是一个干净、最小化的修复,解决了 issue #2573(多卡离线分析时的配置缓存错误)。reset_is_load() 方法被正确地添加到 ProfilerConfig 中,在 analyse_profiling_data 中被正确调用,并且有一个合适的单元测试。未发现任何正确性、安全性、可靠性或破坏性变更的问题。

已审查文件

  • test/profiler/analysis/test_profiler_config.py — 无问题
  • torch_npu/profiler/analysis/_profiler_config.py — 无问题
  • torch_npu/profiler/analysis/_profiling_parser.py — 无问题

⚠️ 已识别出整体风险,但无法提取行内评论,请参考整体评估。

likedislike
chenhao_1209成员
7月3日 评论:

/approve

likedislike
ascend-robotascend-robot成员
7月3日 添加了label:approved
renyujin成员
7月3日 评论:

/lgtm

likedislike
ascend-robotascend-robot成员
7月3日 添加了label:lgtm
ascend-robotascend-robot成员
7月3日 合入了pull request
ascend-robot
ascend-robot成员
7月3日 评论:
流水线 pytorch_gitcode_PR_multiVersion#11800 [ commitID:d7350137 ] 已完成
likedislike