已关闭
[RFC]: Error日志统一增加reason #3951
wanglijun55创建于 8月12日关闭于 20 天前
8月12日 添加了label:rfc
8月12日 关联了看板:FrameworkPTAdapter 版本issue看板
8月12日 添加了label:triage-review
8月12日 修改了issue 的描述
8月12日 将 wanglijun55 设为负责人
8月12日 修改了issue 的描述
8月12日 关联了pull request:[WIP] modify error log
8月12日 添加了label:bot-triaged;删除了label:triage-review
TorchNPU-Bot
8月12日 评论:
8月12日 评论:
检测到当前 issue 已关联 PR,自动添加标签:bot-triaged


9月10日 关联了pull request:fix: append human-readable reason to HCCL/ACL error messages
24 天前 关联了pull request:fix: append human-readable reason to HCCL/ACL error messages
24 天前 关联了pull request:fix: append human-readable reason to HCCL/ACL error messages
24 天前 关联了pull request:fix: append human-readable reason to HCCL/ACL error messages
24 天前 关联了pull request:fix: append human-readable reason to HCCL/ACL error messages
24 天前 关联了pull request:fix: append human-readable reason to HCCL/ACL error messages
20 天前 issue状态由 TODO 改变为 DONE
20 天前 关闭了 issue
20 天前 添加了label:resolved
状态(Status): Draft
作者(Authors): @wanglijun55
创建日期(Created): 2026-08-12
更新日期(Updated): 2026-08-12
相关 Issue/PR: https://gitcode.com/Ascend/pytorch/pull/44441
1. 概述
1.1 简介
用户在排查分布式训练异常时,遇到如下报错:
日志仅打印了 error code
19和 错误码ERR00100,没有说明错误原因和可行的解决,需要额外查阅PTA/CANN的资料来确认,因此提出了日志管理易用性提升的JDC。本 RFC 提出统一的解决方案:为所有errorcode的已有日志,通过检索
error_code_map获取 reason,并将其拼接到已有日志之后。1.2 动机
1.3 目标
"error code is"日志位置完成 reason 补全(排除注释和非日志的 2 处,实际修改 16 处日志语句)。NPUErrorCodes.h中新增HcclErrorCode类(24 个码)和StressDetectErrorCode类(4 个码),与已有的AclErrorCode并列。error_code_map.find()检索 → 将命中的 reason 拼接到日志中。_error_code.py中的正则匹配模式不在范围内)。2. 用例分析
2.1 搜索范围与结果
grep -rn "error code is" --include="*.cpp" --include="*.h" --include="*.py" torch_npu/共命中 18 处。按是否需要修改分类如下:
AclErrorCode类,但日志中未调用 lookupNPUErrorCodes.h:160)和 Python 正则匹配模式(_error_code.py:104),非日志输出2.2 逐文件详情
NPUErrorCodes.hNPUException.hNPU_CHECK_WARN)AclErrorCodeNPUException.hCHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGE)NPUException.hNPU_CHECK_ERROR_CHECK_UCE)AclErrorCodeNPUException.hDEVICE_TASK_ABORT分支)NPUException.hOPS_CHECK_ERROR)AclErrorCodeNPUException.cppAclErrorCodeNPUException.cppAclErrorCodeNPUException.cppAclErrorCodeCalcuOpUtil.hACL_REQUIRE_OK_OP,含 2 路径)AclErrorCodeHCCLUtils.hppHCCL_CHECK_ERROR,含 2 路径)HcclErrorCodeProcessGroupLCCL.cppHcclErrorCodeStress_detect.cppStressDetectErrorCode_error_code.py2.3 场景用例
ACL_REQUIRE_OK_OP宏(所有 op-plugin 算子调用)、NPU_CHECK_ERROR_CHECK_UCE宏(所有 aclrt 调用)HCCL_CHECK_ERROR宏(hcclBroadcast/hcclAllReduce等)、CHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGEProcessGroupLCCL::run_collectiveNPU_CHECK_ERROR_CHECK_UCE的 DEVICE_TASK_ABORT 分支StressDetector::transfer_resultgetDeviceErrorMessage/repair_device_errorOPS_CHECK_ERROR宏(op-plugin 侧)2.4 功能点与质量要求
[Error]: <description>(ACL 风格)或(<description>)(HCCL 风格)。HcclErrorCode/StressDetectErrorCode与已有AclErrorCode保持一致的命名和结构风格。3. 方案设计
3.1 总体方案
核心思路: 为每一处
"error code is"日志引入对应 ErrorCode 类的 lookup 逻辑,实现error code → reason的自动翻译。修改前:
修改后:
3.1.1 关键数据结构:ErrorCode 类
所有错误码映射使用统一的
std::unordered_map<int, std::string>结构,定义在NPUErrorCodes.h中。已有的
AclErrorCode(以 code 100001 为例):torch_npu/csrc/core/npu/NPUErrorCodes.h:namespace c10_npu::acl { class AclErrorCode { public: std::unordered_map<int, std::string> error_code_map = { // ... ~120 entries ... {100001, "ACL uninitialized.\n\ (1)Check whether the acl.init interface has been invoked...\n\ (2)Check whether the initialization interface of the corresponding function..."}, // ... }; };本 RFC 新增的
HcclErrorCode(以 code 19 为例):class HcclErrorCode { public: std::unordered_map<int, std::string> error_code_map = { {1, "parameter error"}, // ... {19, "call network api fail"}, // ← 客户遇到的报错码 // ... {24, "out of memory"}, }; };本 RFC 新增的
StressDetectErrorCode(bitmask 编码):class StressDetectErrorCode { public: std::unordered_map<int, std::string> error_code_map = { {0x1, "bit fail (hardware malfunction)"}, {0x2, "low bit fail (hardware malfunction)"}, {0x4, "high bit fail (hardware malfunction)"}, {0x8, "clear device state fail (voltage recovery failed)"}, }; };3.1.2 关键查找模式:如何在日志位置命中 error_code_map
在每个日志位置,通过以下三步完成 lookup:
// Step 1: 声明 static 局部 ErrorCode 实例(仅首次调用时构造,线程安全) static c10_npu::acl::AclErrorCode err_map; // ACL 错误码 // static c10_npu::acl::HcclErrorCode hccl_err_map; // HCCL 错误码 // static c10_npu::acl::StressDetectErrorCode stress_err_map; // Stress 错误码 // Step 2: 在 error_code_map 中检索 auto it = err_map.error_code_map.find(code); // Step 3: 根据检索结果拼接到日志 // 命中 → 输出 description // 未命中 → 输出 "." 兜底,不破坏原日志格式 const char* reason = (it != err_map.error_code_map.end()) ? it->second.c_str() : ".";在 TORCH_CHECK 宏中的实际写法(ACL 风格,
\n[Error]:前缀):// 修改前(只有数字,没有 reason): TORCH_CHECK((expr) == 0, __func__, ":", __FILE__, ":", __LINE__, " NPU error,NPU error code is:", expr, "\n", c10_npu::acl::AclGetErrMsg(), OPS_ERROR(ErrCode::INTERNAL)); // 修改后(增加 err_map 声明 + lookup + reason 拼接): static c10_npu::acl::AclErrorCode err_map; // ← Step 1: 声明 TORCH_CHECK((expr) == 0, __func__, ":", __FILE__, ":", __LINE__, " NPU error,NPU error code is:", expr, // 原日志不变 (err_map.error_code_map.find(static_cast<int>(expr)) != // ← Step 2: find() 检索 err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[static_cast<int>(expr)] : "."), // ← Step 3: 命中→reason / 未命中→"." "\n", c10_npu::acl::AclGetErrMsg(), OPS_ERROR(ErrCode::INTERNAL));在 HCCL_CHECK_ERROR 宏中的实际写法(HCCL 风格,
(<desc>)后缀):// 修改前: oss << " HCCL function error: " << ... << ", error code is " << Error << ... // 修改后: static c10_npu::acl::HcclErrorCode hccl_err_map; // ← Step 1: 声明 oss << " HCCL function error: " << ... << ", error code is " << Error << (hccl_err_map.error_code_map.find(static_cast<int>(Error)) != // ← Step 2: find() 检索 hccl_err_map.error_code_map.end() ? " " + hccl_err_map.error_code_map[static_cast<int>(Error)] : "") // ← Step 3: 命中→" reason" / 未命中→"" << " " << DIST_ERROR(ErrCode::HCCL) + ".\n";StressDetect 场景(bitmask 需逐位检查):
// Stress 错误码为 bitmask 组合,需要逐位查找 static std::string getStressDetectErrorDesc(int errorCode) { static c10_npu::acl::StressDetectErrorCode err_map; // ← Step 1: 声明 std::string desc; for (const auto& [code, msg] : err_map.error_code_map) { if (errorCode & code) { // ← Step 2: bitmask 匹配 if (!desc.empty()) desc += "; "; desc += msg; // ← Step 3: 拼接所有命中项 } } return desc; // 返回 "" 表示未命中(调用侧兜底为 ".") } // 调用侧: auto desc = getStressDetectErrorDesc(detectResult); ASCEND_LOGW("..., error code is %d.%s", device_id, detectResult, desc.empty() ? "." : ("\n[Error]: " + desc).c_str());3.1.3 实现流程
3.1.4 设计原则
static局部变量,map 仅构造一次,线程安全(只读)。\n[Error]:前缀,HCCL 用(<desc>)后缀),不强行统一。.或""),不会因缺失映射导致崩溃或空输出。3.2 技术选型
if (code==19) desc="call network api fail"std::unordered_map<int, std::string>),在各日志位置声明 static 实例并 lookupAclErrorCode模式一致getErrorDesc(int code)3.3 功能与性能设计
3.3.1 新增 ErrorCode 类
需要在
NPUErrorCodes.h中namespace c10_npu::acl内新增两个 ErrorCode 类:HcclErrorCode(HCCL 通信错误码,24 个):
class HcclErrorCode { public: std::unordered_map<int, std::string> error_code_map = { {1, "parameter error"}, {2, "empty pointer"}, {3, "memory error"}, {4, "internal error"}, {5, "not support feature"}, {6, "not found specific resource"}, {7, "resource unavailable"}, {8, "call system interface error"}, {9, "timeout"}, {10, "open file fail"}, {11, "tcp connect fail"}, {12, "roce connect fail"}, {13, "tcp transfer fail"}, {14, "roce transfer fail"}, {15, "call runtime api fail"}, {16, "call driver api fail"}, {17, "call profiling api fail"}, {18, "call cce api fail"}, {19, "call network api fail"}, {20, "try again"}, {21, "error cqe"}, {22, "error communicator suspending"}, {23, "retry constraint"}, {24, "out of memory"}, }; }; /* hcclError code */StressDetectErrorCode(aclnn 压力检测错误码,4 个):
class StressDetectErrorCode { public: std::unordered_map<int, std::string> error_code_map = { {0x1, "bit fail (hardware malfunction)"}, {0x2, "low bit fail (hardware malfunction)"}, {0x4, "high bit fail (hardware malfunction)"}, {0x8, "clear device state fail (voltage recovery failed)"}, }; }; /* aclnn stress detect error codes */3.3.2 逐位置改动方案
以下按错误码体系分组,说明每处日志的修改方式。
组 A:ACL 错误码(已有
AclErrorCode,新增 lookup)—— 共 12 处这 12 处日志打印的是标准 ACL 错误码,
AclErrorCode类已存在于NPUErrorCodes.h:7-398(含 ~120 个码),只需在日志位置声明 static 实例并追加 lookup 结果。A1.
NPUException.h:35—NPU_CHECK_WARN宏当前代码:
#define NPU_CHECK_WARN(err_code) do { auto Error = err_code; if ((Error) != ACL_ERROR_NONE) { TORCH_NPU_WARN("NPU warning, error code is ", Error, "[Error]: ", (err_map.error_code_map.find(Error) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[Error] : "."), "\n", c10_npu::c10_npu_get_error_message()); } } while (0)→ 已补全,无需修改。
A2.
NPUException.h:173,201—NPU_CHECK_ERROR_CHECK_UCE宏(compact + 非 compact)compact 路径(line 173)和非 compact 主路径(line 201)均已通过
err_map.error_code_map.find()进行 lookup。→ 已补全,无需修改。
⚠️ 但非 compact 的
DEVICE_TASK_ABORT分支(line 195)漏掉了 lookup,见下方 A3。A3.
NPUException.h:195—NPU_CHECK_ERROR_CHECK_UCE的DEVICE_TASK_ABORT分支 🔧调用链路: NPU kernel abort →
NPU_CHECK_ERROR_CHECK_UCE→ DEVICE_TASK_ABORT 分支当前代码:
} else if (error_code == ACL_ERROR_RT_DEVICE_TASK_ABORT) { TORCH_CHECK(false, __func__, ":", __FILE__, ":", __LINE__, " NPU function error: ", (device_error_msg.empty() ? " FORCE STOP" : device_error_msg), ", error code is ", error_code, PTA_ERROR(ErrCode::ACL));修改方案: 宏外层(line 151)已声明
static c10_npu::acl::AclErrorCode err_map;,直接复用。在PTA_ERROR后追加 lookup:} else if (error_code == ACL_ERROR_RT_DEVICE_TASK_ABORT) { TORCH_CHECK(false, __func__, ":", __FILE__, ":", __LINE__, " NPU function error: ", (device_error_msg.empty() ? " FORCE STOP" : device_error_msg), ", error code is ", error_code, PTA_ERROR(ErrCode::ACL), (err_map.error_code_map.find(error_code) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[error_code] : "."));A4.
NPUException.h:230,249—OPS_CHECK_ERROR宏(compact + 非 compact)均已通过
err_map.error_code_map.find()进行 lookup。→ 已补全,无需修改。
A5.
NPUException.cpp:334—checkUceErrAndRepair()当前代码:
static c10_npu::acl::AclErrorCode err_map; err_msg = ... + " NPU error, error code is " + std::to_string(err) + PTA_ERROR(ErrCode::ACL) + (err_map.error_code_map.find(err) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[err] : ".") + ...→ 已补全,无需修改。
A6.
NPUException.cpp:249—getDeviceErrorMessage()🔧调用链路: 异常处理 →
getDeviceErrorMessage()→AclrtGetErrorVerbose自身失败当前代码:
if (ret != ACL_ERROR_NONE) { ASCEND_LOGE("AclrtGetErrorVerbose failed, device is %d, error code is %d.", device, ret); return ""; }修改方案:
if (ret != ACL_ERROR_NONE) { static c10_npu::acl::AclErrorCode err_map; std::string err_info = (err_map.error_code_map.find(ret) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[ret] : "."); ASCEND_LOGE("AclrtGetErrorVerbose failed, device is %d, error code is %d.%s", device, ret, err_info.c_str()); return ""; }A7.
NPUException.cpp:273—repair_device_error()🔧调用链路: UCE 修复 →
repair_device_error()→AclrtRepairError自身失败当前代码:
if (ret != ACL_ERROR_NONE) { ASCEND_LOGE("AclrtRepairError failed, device is %d, error code is %d.", error_info.device, ret); return false; }修改方案: 与 A6 同模式。
if (ret != ACL_ERROR_NONE) { static c10_npu::acl::AclErrorCode err_map; std::string err_info = (err_map.error_code_map.find(ret) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[ret] : "."); ASCEND_LOGE("AclrtRepairError failed, device is %d, error code is %d.%s", error_info.device, ret, err_info.c_str()); return false; }A8.
CalcuOpUtil.h:45,53—ACL_REQUIRE_OK_OP宏 🔧调用链路: 任意 NPU 算子 →
OpParamMaker::InnerRun()/InnerRunOpApi()→ACL_REQUIRE_OK_OP改动要点: 宏内新增
static c10_npu::acl::AclErrorCode err_map;,compact 和非 compact 两路径均追加 lookup 结果。#define ACL_REQUIRE_OK_OP(expr, opstr) do { if (ASCEND_UNLIKELY((expr) != 0)) { std::cout << (opstr) << std::endl; static c10_npu::acl::AclErrorCode err_map; // ← 新增 if (c10_npu::option::OptionsManager::IsCompactErrorOutput()) { std::ostringstream oss; oss << " NPU error,NPU error code is:" << (expr) << (err_map.error_code_map.find(static_cast<int>(expr)) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[static_cast<int>(expr)] : ".") // ← 新增 << "\n" << OPS_ERROR(ErrCode::INTERNAL); ... } else { TORCH_CHECK((expr) == 0, __func__, ":", __FILE__, ":", __LINE__, " NPU error,NPU error code is:", expr, (err_map.error_code_map.find(static_cast<int>(expr)) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[static_cast<int>(expr)] : "."), // ← 新增 "\n", c10_npu::acl::AclGetErrMsg(), OPS_ERROR(ErrCode::INTERNAL)); } } } while (0)组 B:HCCL 错误码(需新增
HcclErrorCode+ 修改日志)—— 共 4 处HCCL 错误码(1-24)尚无对应的 ErrorCode 类。需先在
NPUErrorCodes.h中新增HcclErrorCode类(见 3.3.1),再在以下位置添加 lookup。B1.
HCCLUtils.hpp:26,39—HCCL_CHECK_ERROR宏 🔧调用链路: 分布式通信 →
hcclBroadcast/hcclAllReduce等 →HCCL_CHECK_ERROR修改方案(compact 路径 + 非 compact 路径):
#define HCCL_CHECK_ERROR(err_code, ...) do { auto Error = err_code; if ((Error) != HCCL_SUCCESS) { CHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGE(Error); static c10_npu::acl::HcclErrorCode hccl_err_map; // ← 新增 if (c10_npu::option::OptionsManager::IsCompactErrorOutput()) { std::ostringstream oss; oss << " HCCL function error: " << getErrorFunction(#err_code, ##__VA_ARGS__) << ", error code is " << Error << (hccl_err_map.error_code_map.find(static_cast<int>(Error)) != hccl_err_map.error_code_map.end() ? " " + hccl_err_map.error_code_map[static_cast<int>(Error)] : "") // ← 新增 << " " << DIST_ERROR(ErrCode::HCCL) + ".\n"; ... } else { auto retmsg = std::string(__func__) + ":" + __FILE__ + ":" + std::to_string(__LINE__) + " HCCL function error: " + getErrorFunction(#err_code, ##__VA_ARGS__) + ", error code is " + std::to_string(Error) + (hccl_err_map.error_code_map.find(static_cast<int>(Error)) != hccl_err_map.error_code_map.end() ? " " + hccl_err_map.error_code_map[static_cast<int>(Error)] : "") + // ← 新增 " " + DIST_ERROR(ErrCode::HCCL) + ".\n" + c10_npu::c10_npu_get_error_message(); ... } } } while (0)B2.
NPUException.h:140—CHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGE宏 🔧调用链路:
HCCL_CHECK_ERROR→CHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGE(HCCL 错误码在此宏中需要 HcclErrorCode 查找)当前此宏已有 AclErrorCode lookup,但缺少 HcclErrorCode lookup。需增加:
#define CHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGE(err_code) ... static c10_npu::acl::AclErrorCode err_map; static c10_npu::acl::HcclErrorCode hccl_err_map; // ← 新增 TORCH_CHECK(false, __func__, ":", __FILE__, ":", __LINE__, " NPU function error: ", device_error_msg, ", error code is ", error_code, (err_map.error_code_map.find(error_code) != err_map.error_code_map.end() ? "\n[Error]: " + err_map.error_code_map[error_code] : (hccl_err_map.error_code_map.find(error_code) != hccl_err_map.error_code_map.end() ? "\n[Error]: " + hccl_err_map.error_code_map[error_code] : ".")), // ← 新增 HCCL fallback PTA_ERROR(ErrCode::ACL));B3.
ProcessGroupLCCL.cpp:254— LCCL 操作错误 🔧调用链路: LCCL 通信 →
ProcessGroupLCCL::run_collective→fn()返回非零LCCL 底层复用 HCCL 错误码体系,使用同一个
HcclErrorCode类。需新增 include。修改方案:
#include "torch_npu/csrc/core/npu/NPUErrorCodes.h" // ← 新增 // line 254: auto ret = fn(inputs[i], outputs[i], lcclComms[i], lcclStream); static c10_npu::acl::HcclErrorCode hccl_err_map; // ← 新增 TORCH_CHECK(ret == 0, "LCCL function error:", opTypeToString(opType).c_str(), ", error code is ", ret, (hccl_err_map.error_code_map.find(ret) != hccl_err_map.error_code_map.end() ? " (" + hccl_err_map.error_code_map[ret] + ")" : ""), // ← 新增 "\n"); // line 191 同步修复("error code:" 而非 "error code is",同类问题): TORCH_CHECK(ret == 0, "init lccl comm failed, error code: ", ret, (hccl_err_map.error_code_map.find(ret) != hccl_err_map.error_code_map.end() ? " (" + hccl_err_map.error_code_map[ret] + ")" : ""), PTA_ERROR(ErrCode::INTERNAL));组 C:StressDetect 错误码(需新增
StressDetectErrorCode+ 修改日志)—— 共 4 处C1-C4.
Stress_detect.cpp:103,104,112,113—transfer_result()函数 🔧调用链路:
torch_npu.npu.stress_detect(detect_type='aic')→_npu_stress_detect()→StressDetector::perform_stress_detect()→transfer_result()修改方案: StressDetect 错误码为 bitmask,需逐位检查。先在
NPUErrorCodes.h新增StressDetectErrorCode类(见 3.3.1),再在Stress_detect.cpp中新增辅助函数并修改 4 处日志 + 1 处硬编码。#include "torch_npu/csrc/core/npu/NPUErrorCodes.h" // ← 确认 include // 新增辅助函数 static std::string getStressDetectErrorDesc(int errorCode) { static c10_npu::acl::StressDetectErrorCode err_map; std::string desc; for (const auto& [code, msg] : err_map.error_code_map) { if (errorCode & code) { if (!desc.empty()) desc += "; "; desc += msg; } } return desc; } // transfer_result() 修改后: int StressDetector::transfer_result(int detectResult) { int ret = kDetectFailed; switch (detectResult) { case 0: ret = kDetectSucceeded; ASCEND_LOGI("..., device id is %d.", device_id); break; case ACLNN_STRESS_BIT_FAIL: case ACLNN_STRESS_LOW_BIT_FAIL: case ACLNN_STRESS_HIGH_BIT_FAIL: ret = kDetectFailedWithHardwareFailure; { auto hw_desc = getStressDetectErrorDesc(detectResult); // ← 新增 ASCEND_LOGW("..., device id is %d, error code is %d.%s", device_id, detectResult, // ← 修改 hw_desc.empty() ? "." : ("\n[Error]: " + hw_desc).c_str()); // ← 新增 TORCH_NPU_WARN("..., device id is ", device_id, ", error code is ", detectResult, // ← 修改 hw_desc.empty() ? "." : "\n[Error]: " + hw_desc); // ← 新增 } break; case ACLNN_CLEAR_DEVICE_STATE_FAIL: { auto fail_desc = getStressDetectErrorDesc(detectResult); // ← 新增(替代原硬编码) ASCEND_LOGW("..., device id is %d, error code is %d.%s", device_id, detectResult, // ← 修改 fail_desc.empty() ? "." : ("\n[Error]: " + fail_desc).c_str()); // ← 新增 TORCH_CHECK(false, "..., error code is ", detectResult, // ← 修改 fail_desc.empty() ? "." : "\n[Error]: " + fail_desc, // ← 新增 PTA_ERROR(ErrCode::ACL)); } break; default: ret = kDetectFailed; { auto fail_desc = getStressDetectErrorDesc(detectResult); // ← 新增 ASCEND_LOGW("..., device id is %d, error code is %d.%s", device_id, detectResult, // ← 修改 fail_desc.empty() ? "." : ("\n[Error]: " + fail_desc).c_str()); // ← 新增 TORCH_NPU_WARN("..., device id is ", device_id, ", error code is ", detectResult, // ← 修改 fail_desc.empty() ? "." : "\n[Error]: " + fail_desc); // ← 新增 } break; } return ret; }3.4 安全隐私与DFX设计
error code is <数字>模式,不受影响(数字仍在原位置)。对于依赖整行正则匹配的工具,需确认追加的[Error]: <描述>或(<描述>)不会导致误匹配。HcclErrorCode/StressDetectErrorCode与已有AclErrorCode同文件(NPUErrorCodes.h)、同命名空间(c10_npu::acl),新成员加入时只需在对应 map 中追加条目。static局部变量,仅构造一次,线程安全(只读)。map 未命中时输出.或空字符串,不会因缺失映射导致崩溃。3.5 编程与调用设计
本提案不涉及对外 API 变更,全部改动在已有宏/函数内部实现,对外接口透明。
3.5.1 受影响模块
NPUErrorCodes.hHcclErrorCode(24 码)+StressDetectErrorCode(4 码)HCCLUtils.hppHCCL_CHECK_ERROR宏内新增 HcclErrorCode lookupNPUException.hCHECK_AND_THROW_ERROR_WITH_SPECIFIC_MESSAGE新增 HcclErrorCode fallback;DEVICE_TASK_ABORT 分支补 lookupNPUException.cppProcessGroupLCCL.cppCalcuOpUtil.hACL_REQUIRE_OK_OP宏内新增 AclErrorCode lookupStress_detect.cpp3.5.2 ErrorCode 类总览(变更后)
AclErrorCodeNPUErrorCodes.h:7-398HcclErrorCodeNPUErrorCodes.hStressDetectErrorCodeNPUErrorCodes.h4. 测试设计
4.1 可复用的已有测试用例
ACL_REQUIRE_OK_OP(CalcuOpUtil.h)third_party/op-plugin/test/test_base_ops/test_npu_scaled_mm.py—test_npu_scaled_mm_invalid_*系列assertRaisesRegex检查[Error]NPU_CHECK_ERROR_CHECK_UCE(NPUException.h)torch_npu.npu.set_device(-1)触发异常[Error]: Invalid device.HCCL_CHECK_ERROR(HCCLUtils.hpp)test/npu/test_c10d.py)覆盖正常通信路径;异常路径需构造DEVICE_TASK_ABORT(NPUException.h)test/npu/test_uce.py—monitor()捕获 "FORCE STOP"[Error]: The aicpu execution is abnormal.4.2 建议新增的测试用例
test/npu/test_stress_detect.pytorch_npu.npu.stress_detect('aic'),验证返回码及日志test/npu/test_error_message.py[Error]description 或(<description>)test/npu/test_lccl.py4.3 批量验证脚本
#!/bin/bash set -e echo "=== 1. ACL_REQUIRE_OK_OP (CalcuOpUtil.h) ===" # compact 模式 TORCH_NPU_COMPACT_ERROR_OUTPUT=1 python -c " import torch, torch_npu try: torch_npu.npu.scaled_mm(torch.randn(2,3).npu(), torch.randn(3,2).npu(), scale_a=torch.randn(1).npu().to(torch.int32)) except RuntimeError as e: assert '[Error]' in str(e), 'FAIL: compact mode missing [Error]' print('PASS: compact mode') " # 非 compact 模式 python -c " import torch, torch_npu try: torch_npu.npu.scaled_mm(torch.randn(2,3).npu(), torch.randn(3,2).npu(), scale_a=torch.randn(1).npu().to(torch.int32)) except RuntimeError as e: assert '[Error]' in str(e), 'FAIL: non-compact mode missing [Error]' print('PASS: non-compact mode') " echo "=== 2. NPU_CHECK_ERROR (NPUException.h) ===" python -c " import torch, torch_npu try: torch_npu.npu.set_device(-1) except RuntimeError as e: assert '[Error]' in str(e), 'FAIL: NPU_CHECK_ERROR path missing [Error]' print('PASS: NPU_CHECK_ERROR path') " echo "=== 3. HCCL_CHECK_ERROR (HCCLUtils.hpp) ===" # 需要多卡环境 # python -c " # import os; os.environ['MASTER_ADDR']='127.0.0.1'; os.environ['MASTER_PORT']='29500' # import torch; import torch_npu # torch.distributed.init_process_group(backend='hccl', rank=0, world_size=1) # torch.distributed.barrier() # " echo "=== 4. ASCEND 日志检查 ===" grep -r "\[Error\]:" /var/log/npu/ascend_log/device-*/device-*.log 2>/dev/null | tail -5 || \ echo "WARN: No device log found."4.4 验证检查清单
5. 缺点和风险
TORCH_NPU_COMPACT_ERROR_OUTPUT=1控制error code is <数字>正则的脚本可能受追加文本干扰Breaking Change: 无。所有改动在已有日志输出基础上追加字段,不影响 API 签名或行为。
6. 现有技术
本方案参考了
NPUErrorCodes.h中已有的AclErrorCode类(std::unordered_map<int, std::string>存储 ~120 个错误码与描述映射)的设计模式。AclErrorCode已在NPU_CHECK_WARN(NPUException.h:33)、NPU_CHECK_ERROR_CHECK_UCE(NPUException.h:151)、OPS_CHECK_ERROR(NPUException.h:223)、checkUceErrAndRepair(NPUException.cpp:332)等位置被正确使用并验证有效。本 RFC 将该模式系统化推广到所有
"error code is"日志位置,并扩展覆盖 HCCL 和 StressDetect 两个新的错误码体系。7. 未解决问题
ACLNN_CLEAR_DEVICE_STATE_FAIL分支的硬编码推断("Voltage recovery failed"对应 0x8)。需与 CANN 文档/HW 团队确认ACLNN_STRESS_BIT_FAIL(0x1) /ACLNN_STRESS_LOW_BIT_FAIL(0x2) /ACLNN_STRESS_HIGH_BIT_FAIL(0x4) 的官方描述是否准确。HcclErrorCode,码值 1-24)。需确认 LCCL 是否有独立的错误码空间(超出 1-24 范围的码)。AclErrorCode路径使用\n[Error]: <description>格式,HcclErrorCode路径使用(<description>)后缀格式。RFC 通过后是否需统一为一种格式?建议保持现有差异(各体系历史原因),或在后续 RFC 中统一。附录
NPUErrorCodes.h中现有AclErrorCode实现欢迎加入社区,感谢您对社区的贡献 🎉!