已合并
[docs] Support pinned_max_round_threshold_mb and pinned_max_cached_size_mb #41563
liujunzhu创建于 7月13日
[docs] Support pinned_max_round_threshold_mb and pinned_max_cached_size_mb #41563
已合并
Pull Request已成功合入, 合并人@ascend-robot
(感谢 liujunzhu 的贡献)7月13日 关联了issue:[Feature]: 支持pinned_max_round_threshold_mb和pinned_max_cached_size_mb内存优化项
atomgit-bot
7月13日 评论:
7月13日 评论:
ascend-robot
7月13日 评论:
7月13日 评论:
atomgit-bot
7月13日 评论:
7月13日 评论:
代码审查
审查总结
对唯一变更文件 docs/zh/environment_variable_reference/PYTORCH_NPU_ALLOC_CONF.md 的审查已完成。
发现统计:
- P0: 0
- P1: 0
- P2: 0
- P3: 1(
pinned_max_round_threshold_mb和pinned_max_cached_size_mb缺少版本要求说明)
各文件审查结论:
docs/zh/environment_variable_reference/PYTORCH_NPU_ALLOC_CONF.md:新增的两个选项文档内容与 PR 设计方案一致,描述准确,示例无问题,交互约束说明正确。唯一不足是未标注版本要求。
整体风险评估: 低风险。这是纯文档变更,不涉及代码逻辑修改。文档内容事实正确,唯一的小缺陷是版本要求遗漏,但不影响正确使用——用户在低版本上使用时会收到框架已有的 "not support key" 警告,不会导致崩溃或数据损坏。
⚠️ 已识别出整体风险,但无法提取行内评论,请参考整体评估。


不准确?
此处折叠了69条消息 查看更多
7月15日 添加了label:ci-pipeline-passed
ascend-robot
7月15日 评论:
7月15日 评论:
流水线 PR-pipeline_pytorch#45677 [ commitID:b4a1760f ] 已完成
>>>代码风格自动修复执行成功(无修复内容)
| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | 🛑 | >>> |
| Build_ARM | 🛑 | >>> | |
| Build_LibTorch_x86 | 🛑 | >>> | |
| Build_LibTorch_ARM | 🛑 | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | CodeCheck | ✅ | >>> |
| check_error | ✅ | >>> | |
| CodeCheck_lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_X86_Part_01 | 🛑 | >>> |
| UT_X86_Part_02 | 🛑 | >>> | |
| UT_ARM_A3_Part_01 | 🛑 | >>> | |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | 🛑 | >>> | |
| UT_ARM_A2_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_02 | 🛑 | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


7月15日 关闭了关联的issue
7月15日 合入了pull request
ascend-robot
7月15日 评论:
7月15日 评论:
流水线 pytorch_gitcode_PR_multiVersion#12613 [ commitID:b4a1760f ] 已完成


【合入来源】
【修改方案】
PyTorch 社区在
AcceleratorAllocatorConfig中新增了pinned_max_round_threshold_mb和pinned_max_cached_size_mb两个 pinned memory 分配器配置选项,用于缓解大块 pinned memory 的内存浪费问题:pinned_max_round_threshold_mb:分配尺寸向上取整到 2 的幂次的上限阈值(MB)。超过此阈值的分配使用精确请求大小,跳过 power-of-2 取整。默认size_t::max()(即禁用)。pinned_max_cached_size_mb:缓存到 free list 的块大小上限阈值(MB)。超过此阈值的块在释放时立即归还给 OS,不再进入 free list 缓存。默认size_t::max()。torch_npu 的
NPUCachingHostAllocatorImpl继承自 PyTorch 的CachingHostAllocatorImpl,基类allocate()与maybe_cache_block()中已实现这两个阈值的核心逻辑(通过虚方法读取配置)。torch_npu 需在配置层与运行时层完成对接,使社区新增选项能在 NPU pinned memory 路径生效。【资料变更】
是,环境变量
PYTORCH_NPU_ALLOC_CONF新增pinned_max_round_threshold_mb和pinned_max_cached_size_mb配置项。【接口变更】
环境变量
PYTORCH_NPU_ALLOC_CONF新增pinned_max_round_threshold_mb和pinned_max_cached_size_mb配置项。【功能验证】
本PR仅修改资料,不涉及功能验证。
【CheckList】