当前Pull Request已关闭, 关闭人@FengHaozhan
变更摘要
本 PR 为 arch5162 平台的运行时线程 AICPU 引入 DataDump 能力。变更围绕一条完整的 DataDump 链路展开:扩展插件协议接口(新增 DataDump SQE 子类型、状态码与请求结构),新增 datadump 子模块负责 TLV 解析(DataDumpParser)、模型/算子信息管理(DataDumpManager)与统计/张量文件落盘(DataDumpWriter),并在 RuntimeThreadAicpuService 中接入 worker 启动、DataDump 加载与回调报告分流处理;同时调整了 API 层 DatadumpInfoLoad 的归档与 arch5162 下发路径,以及在 AI Core SQE 中打上 DataDump 标志。
主要改动
-
插件协议接口扩展:
runtime_thread_aicpu_plugin.h新增RuntimeThreadAicpuSqeSubtype::DATADUMP、7 个DATADUMP_*状态码,以及RuntimeThreadAicpuStartRequest、RuntimeThreadAicpuDumpInfoRequest结构与startWorker、loadDumpInfo两个函数指针,插件 API 契约随之扩充。 -
新增 DataDump 解析、管理与落盘模块: 新增
data_dump_parser、data_dump_manager、data_dump_writer、data_dump_types.hpp、data_dump_tlv.hpp,实现 TLV 二进制协议解析(含版本/魔数校验与层级去重)、按TaskKey索引模型与算子信息、按数据类型生成统计 CSV 或按 TLV 布局写出张量文件,并对文件/目录创建与写入失败返回对应状态。 -
服务层接入与回调分流:
RuntimeThreadAicpuService新增StartWorker、LoadDumpInfo,将EnsureStarted参数由RuntimeThreadAicpuKernelRequest改为deviceId/tsId,并把原先的ProcessOneReport/FinishReport拆分为ProcessAicpuReport/ProcessDumpReport与带RuntimeThreadAicpuSqeSubtype的FinishReport,在ProcessReports中按报告reserved字段路由 AICPU 与 DATADUMP 两类回调。 -
运行时 API 层下发路径调整: 将
DatadumpInfoLoad从api_impl.cc移至api_impl_cpu_kernel.cc,并在api_impl_arch5162.cc新增实现(校验flag、上下文与设备后调用LoadRuntimeThreadAicpuDumpInfo);runtime_thread_aicpu.cc增加LoadDumpInfo转发与插件表校验,runtime_thread_aicpu.hpp导出对应接口。 -
AI Core SQE 打标:
davinci_task_arch5162.cc在ConstructAICoreSqeForDavinciTask中当kernelFlag含RT_KERNEL_DUMPFLAG时设置header.postP与res7[1] |= SQE_BIZ_FLAG_DATADUMP,并加入偏移静态断言约束标志落在第 13 个 word。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.
PR Approval Progress
⚠️ This PR does not yet meet the following requirements:lgtm (requires ≥ 2 person(s) per module)、approve (requires ≥ 1 person(s) per module)
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| pkg_inc | ❌ (0/2)(You can also ask: yanmingxiang, Reyn52166, 王涛, 张子菁, 卢煜坤) | ❌ (0/1)(You can also ask: 张子菁, 卢煜坤, 侯延保, 王涛) |
| src/aicpu_sched | ❌ (0/2)(You can also ask: 张子菁, 卢煜坤, Reyn52166, LiWei79, yanmingxiang) | ❌ (0/1)(You can also ask: LiWei79, ZhaiPeiChao) |
| src/runtime | ❌ (0/2)(You can also ask: sunnana_004434229, yanmingxiang, weiying_101, WangHuiH, zhangpengpeng8) | ❌ (0/1)(You can also ask: Turing_JasonWen, WangHuiH, YzQnWyx, turing_yhy, Henry_ascend) |
| tests/ut/aicpu_sched | ❌ (0/2)(You can also ask: LiWei79, 侯延保, zhangpengpeng8, Reyn52166, 张子菁) | ❌ (0/1)(You can also ask: ZhaiPeiChao, LiWei79) |
| tests/ut/runtime | ❌ (0/2)(You can also ask: Reyn52166, 张子菁, turing_yhy, zhangpengpeng8, zhaozhixuan) | ❌ (0/1)(You can also ask: sunnana_004434229, 卢煜坤, Axiaolei, turing_yhy, WangHuiH) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
FengHaozhan, thanks for your pull request. All authors of the commits have signed the CLA. 👍


🔴 Critical:这里把现有 rtDatadumpInfoLoad 载荷直接转给 Host 侧 TLV 解析,破坏了当前跨仓接口契约。
当前 GE V1/OM2 在调用该接口前都会先把 OpMappingInfo protobuf H2D 拷到 Device,再传入 Device 地址(runtime/v1/common/dump/data_dumper.cc:1016-1023、runtime/om2/dump/data_dump_impl.cc:128-140);而新路径最终在 DataDumpParser::Parse 中直接对 dumpInfo 做 Host memcpy,并要求内容是 0x5A5A5A5A TLV。这样现有 arch5162 调用会把 Device 地址当 Host 指针解引用,可能直接异常,即使地址可访问也必然因 protobuf/TLV 格式不符而加载失败。
建议先同步生产端与接口契约:要么由 GE 通过明确的 Host-TLV 接口传入完整 TLV,要么保留现有 Device-protobuf 通路;同时补一条从真实 GE 调用形态进入该实现的集成用例。


modelName 直接 JoinPath 进 dump 目录路径且不经 SanitizeFileNamePart(opName/opType 有消毒、modelName 没有);parser 对 modelName 无字符集校验。TLV 中 modelName 含 "../" 或绝对路径分量时,CreateDirectories 按 '
mkdir 可在 dumpPath 之外建目录/写文件。
建议 parser 或 writer 侧对 modelName 做与 opName 相同的消毒。


modelName不需要,GE能保证模型名没有这些特殊字符
changed this line on 9edb0d0e view diff detail
changed this line on 9edb0d0e view diff detail
未知 CallbackReportType(reserved 非 0/1):
(1) 不调 FinishReport → SQ finish command 不回写 → 该 report 对应 event 永不 complete,上层 stream/event 同步挂起;
(2) success=false → WorkerLoop 置 failed_ 并永久退出 → 整个 runtime_thread_aicpu 服务(含普通 AICPU 算子)不可用。
既有失败策略(funcPtr 校验失败等)为"单条失败→setStreamError+回写 finish,不杀 worker",未知 type 处理与之不一致。
请确认是否存在上述问题,请确认是否OK?


changed this line on 9edb0d0e view diff detail
ReadString/ParseStringList 对 TLV 长度无上限:恶意/损坏 TLV(length≈4GB)触发数 GB 单次分配,bad_alloc 兜底不崩溃但内存尖峰可能 OOM-kill 调度进程。
建议增加单字段上限(如 1MB)或累计限流。


tlv不来自外部输入,而是内部组件GE提供的,可以保证没有超长tlv存在
changed this line on 9edb0d0e view diff detail
changed this line on 9edb0d0e view diff detail
changed this line on 9edb0d0e view diff detail
compile


| 🚀 CI 流水线已启动 |
|---|
| 📋 执行详情: 点击查看流水线 |


Pull Request
描述
请清晰准确地描述本次 Pull Request 的意图和变更内容。
变更类型
请选择本次引入的变更类型:
关联的Issue
如何测试
描述测试此变更的步骤和前提条件:
1.
2.
核对清单
其他信息
在此添加任何其他关于本次 PR 的说明。