已开启
[Feature]: 统一 symmetric memory、multicore 与 custom ops 的编译打包流程 #333
hedongdong创建于 8月14日
8月14日 关联了pull request:refactor: unify optional native build and packaging
8月14日 修改了issue 的描述
8月14日 修改了issue 的描述
8月14日 修改了issue 的描述
8月18日 修改了issue 的描述
8月25日 修改了issue 的描述
8月25日 修改了issue 的描述
8月25日 修改了issue 的描述
8月25日 修改了issue 的描述
8月25日 修改了issue 的描述
hedongdong
8月25日 评论:
8月25日 评论:
使用 libhyper_parallel_shmem*.so 而不是原生 libshmem*.so,核心原因是进程内动态库隔离,不是修改 SHMEM 功能。
HyperParallel 的 wheel 会自带一套锁定版本的 SHMEM SDK/runtime。如果仍使用原生 SONAME:
- 进程可能已被 CANN、torch_npu 或其他组件加载另一份
libshmem.so。 - ELF 动态链接器可能按相同 SONAME 复用已加载库,导致 HyperParallel 实际绑定到错误版本。
LD_LIBRARY_PATH顺序也可能让 HyperParallel 加载到系统 SHMEM,而不是 wheel 内经过验证的版本。- MindSpore 用户会被迫额外安装并维护外部 SHMEM,wheel 不再相对自包含。
因此采用私有名称:
libshmem.so → libhyper_parallel_shmem.so
libshmem_utils.so → libhyper_parallel_shmem_utils.so
并让 HyperParallel 自己的 ELF 通过相对 DT_RUNPATH=$ORIGIN/... 解析这些私有依赖。这样系统 SHMEM 与 HyperParallel SHMEM 可以在同一进程共存,且不会依赖全局库搜索顺序。
外部 wrapper 的作用只是控制输出文件名、SONAME、链接关系和 RUNPATH;SHMEM 源码仍来自锁定的上游版本,没有直接修改上游源码,也没有形成业务功能分叉。
只有一种情况下可以不用私有 SONAME:HyperParallel 不再随 wheel 携带 SHMEM,明确要求所有用户预装唯一且兼容的 SHMEM runtime/SDK。这样能直接依赖原生 libshmem.so,但依赖安装、版本匹配和冲突诊断都转移给用户,MindSpore 场景的易用性会明显下降。


8月28日 关联了pull request:fix: accept CANN 9.1.0 and newer for native builds
8月28日 关联了pull request:test: move optional native cases to Level 1
8月29日 关联了pull request:refactor: decouple multicore build from CANN source dependencies
28 天前 关联了pull request:fix(native): use public framework build interfaces
25 天前 关联了pull request:refactor(native): decouple framework adapter builds
背景与原有问题
HyperParallel 的 symmetric memory(单边通信)、multicore/HyperMegaMoe 和 MindSpore custom ops 同时包含 C++、AscendC、CANN vendor 和 CPython 扩展。原流程可支持早期开发,但存在以下发布与维护问题:
build.sh -> setup.py::BuildPy.run()隐式调用三个组件脚本,bdist_wheel同时承担 CANN 检查、依赖准备、native 编译和 Python 打包。直接执行setup.py bdist_wheel也会触发 native 编译,形成显式/隐式双入口,失败时可能只留下 setup warning。mega_moe.tar.gz无法由当前提交、上游 commit、CANN 和 SoC 参数确定性重建,也不适合作为 910B/910C 及后续硬件扩展基础。MegaMoe/MegaMoeGrad/aclnnMegaMoe*与 CANN 9.1 内置算子重名;前向和反向分别携带同 SONAME 的libcust_opapi.so并全局加载,存在错误绑定风险。.so或依赖带 torch_npu 的 Python wheel。build/lib和预编译目录;CANN custom OPP 不能由PYTHONPATH代替,框架导入后再修改 OPP/动态库环境已经太晚。requirements.txt;包含 CPython module 的 adapter 必须按 cp310/cp311/cp312 分别出包;不同 host arch 不应混包。目标方案
build.sh是唯一推荐的全量入口,显式完成环境检查、依赖准备、组件编译、payload 组装和 wheel 打包;setup.py只打包已有 payload,不下载依赖、不编译 native。ascend910b(910B)和ascend910_93(910C)kernel。libhyper_parallel_shmem*.so私有 SONAME 和相对 RUNPATH;multicore 与 symmetric memory 共用完整 SHMEM SDK。HyperMegaMoe/HyperMegaMoeGrad、aclnnHyperMegaMoe*;前后向进入同一 vendor 和一个libcust_opapi.so,公开头文件使用独立 include guard。ascend910b(910B)和ascend910_93(910C)分别生成 kernel/config;统一 vendor 固定选择 910C 优先的 canonical host,校验共同构建输入身份和动态 ABI,丢弃其他 SoC 的重复 host 变体,不对独立链接 ELF 做不可靠的“语义等价”推断。build/native/payload/hyper_parallel;multicore 在框架进程启动前显式 source OPP 脚本,不写~/.bashrc、不安装.pth、不在 Python import 后修改LD_LIBRARY_PATH。--strict on要求所选组件全部成功。正式版本发布由版本构建、Level1/全量用例和人工评审决定。构建流程:原有与当前
对使用者而言,完整构建命令仍是
./build.sh;变化在于职责从 setuptools 隐式回调改为顶层显式编排。原有流程:
当前流程:
三个组件脚本仍可独立用于局部开发并刷新各自 payload slice;重型 SHMEM/per-SoC vendor 缓存默认复用,轻量 framework adapter 每次从按框架身份隔离的干净目录重建。
--clean只清理所选组件 work/install,不删除下载依赖。用户入口
最简全量构建:
默认尝试 multicore=
all、shmem=all、custom-ops=on、SoC=ascend910b,ascend910_93,生成 PYTHONPATH payload 和当前 CPython/host arch wheel。默认strict=off,所以“wheel 已生成”不代表所有 optional 组件均成功;要求完整成功时使用:统一入口的完整参数如下:
--multicorealloff、mindspore/ms、torch/pytorch、all/both--shmemalloff、mindspore/ms、torch/pytorch、all/both--custom-opsonon、off--soc-listascend910b,ascend910_93ascend910b、ascend910_93、ascend950的逗号分隔组合ascend910b(910B)和ascend910_93(910C)已实现,ascend950当前产生明确的 optional failure--strictoffon、offoff时删除失败组件的半成品、保留其他成功组件并继续出 wheel;on时任一选中组件失败即终止,不出 wheel--jobsnproc--cleanbuild/native/deps下载缓存-h/--help接口采用以下固定约定:
build/native/payload/hyper_parallel,并使用当前 shell 中的 Python 生成对应 CPython ABI、当前 host arch 的一个 wheel;不提供--python和跨架构交叉打包参数。--wheel开关;执行成功后同时输出 payload 路径和本次 wheel 的精确路径。--prepare-deps、--offline、--deps-dir等常规用户参数。--shmem off不能关闭 multicore 必需的 symmetric memory。例如--multicore torch --shmem off的最终 symmetric-memory target 仍为torch;--multicore torch --shmem mindspore的最终 target 为all。strict=off只保证核心 HyperParallel wheel 能输出;是否包含某个 optional native 组件以独立日志和 wheel 内容为准。正式发布仍由 Level1/全量用例和人工流程决定。常用构建组合:
# 默认全量:双框架 symmetric memory、双框架 multicore、MindSpore custom ops、ascend910b+ascend910_93。 ./build.sh # MindSpore native 组件。 ./build.sh --multicore mindspore --shmem mindspore --custom-ops on # Torch native 组件,不构建 MindSpore custom ops。 ./build.sh --multicore torch --shmem torch --custom-ops off # 仅核心 Python wheel,不构建 optional native 组件。 ./build.sh --multicore off --shmem off --custom-ops off # 仅携带 910C 目标,并要求全部选中组件成功。 ./build.sh --soc-list ascend910_93 --strict on # 清理选中组件缓存并指定并行度。 ./build.sh --clean --jobs 32构建机必须预装 CANN 9.1.0。
ASCEND_HOME_PATH未设置且/usr/local/Ascend/cann/set_env.sh存在时,build.sh只在自身子进程中 source;自定义 CANN 路径由调用方预先 source,不修改父 shell。wheel 安装态运行 multicore:
pip install /exact/path/hyper_parallel-0.1.0-cpXY-cpXY-linux_aarch64.whl source /path/to/CANN-9.1.0/set_env.sh source "$(command -v hyper_parallel_multicore_set_env.bash)" python application.pyPYTHONPATH 开发态:
source /path/to/CANN-9.1.0/set_env.sh ./build.sh export PYTHONPATH=/path/to/hyper-parallel:${PYTHONPATH:-} source build/native/payload/hyper_parallel/core/multicore/lib/set_env.bash python application.py激活脚本必须早于 MindSpore、torch 或 torch_npu 导入。未激活、激活过晚、wheel 缺 payload、adapter/依赖加载失败分别返回稳定
HP-NATIVE-*reason code 和精确恢复命令。wheel 安装只安装可 source 的 locator,不自动执行,不修改用户持久环境;pip uninstall按 wheelRECORD删除 locator。当前实现与支持边界
>=2.10;Torch/torch_npu 沿用 HyperParallel 既有 extras 与 CANN 配套,不在 native lock 中固定版本。已完成验证与剩余矩阵
./build.sh --strict on全量成功,生成包含ascend910b/ascend910_93kernel、双框架 adapter、symmetric memory 和 custom ops 的单一 wheel。aclnnHyperMegaMoe*符号、ELF arch、RUNPATH、私有 SHMEM SONAME/NEEDED、安装 locator 和RECORD已完成静态核验。35 passed;真实双 SoC vendor 合并确认选择ascend910_93host 并同时保留 910B/910C kernel;Lizard、Shell/Python 语法及 diff 检查通过。本 ISSUE 不新增 Python 业务 API;Python 仍使用
mega_moe/mega_moe_grad及现有 symmetric memory/custom ops API。