已开启
[RFC]: 删除platform #377
yangzhenzhang创建于  19 天前
yangzhenzhang成员
19 天前 创建

动机(Motivation).

1. 背景与现状

hyper_parallel/platform/ 当前的架构是 双实现 + 单一分发器:

  • platform/platform.py(约 67KB)定义抽象的 Platform 基类,契约面高达 100+ 个属性(collective、differentiable_* 异步族、tensor 工厂、checkpoint/swap、stream/random、param 自省、custom_ops 等)。
  • platform/mindspore/platform.py(约 83KB)与 platform/torch/platform.py(约 64KB)是两套各自独立、平行实现的具体平台类。
  • 分发逻辑 get_platform()(platform/platform.py:99-126)返回全局单例,默认 MindSpore,可被环境变量 HYPER_PARALLEL_PLATFORM 覆盖,失败时 ImportError 回退 torch:
# platform.py
try:
    return get_mindspore_platform()
except ImportError:
    return get_torch_platform()
  • 整个代码库(core/、trainer/、models/、collectives/、integration/、dmodule/)几乎全部通过 from hyper_parallel import get_platform 拿到单例,再在模块顶层绑定 Tensor = platform.Tensor、platform.get_rank() 等使用。core/ 层不直接 import torch 或 import mindspore,而是通过接缝与框架解耦。

2. 本次任务目标

删除platform,代码直接使用torch进行实现。

目标设计.

删除platform,代码直接使用torch进行实现。

需要设计评审DFX建议.

NA

相关的RFCs和API

NA

完整的反馈期限.

CC List.

其他补充说明.

Thanks for contributing 🎉! The hyper-parallel core team hosts a biweekly RFC review session, while most RFCs can be discussed online, you can optionally sign up for a slot to discuss your RFC online.

Before submitting a new issue...

likedislike
Yyangzhenzhang成员
19 天前 添加了label:RFC
hedongdong成员
19 天前 评论:

Custom_ops

一、任务边界

ISSUE #377 的总体目标是删除整个 Platform 层。本方案对应其中一个独立子任务:

完整删除 HyperParallel 自有的 legacy custom_ops 特性,包括 API、MindSpore 实现、Platform 接缝、编译入口、分布式规则、资料和测试。

本任务不要求全仓 custom_ops 字符串清零,重点是确保 legacy custom_ops 特性及其能力入口完整消失。

前置依赖:KDA/GDN 迁移

hyper_parallel/platform/torch/custom_ops/kda/ 和 gdn/ 下是实际工作的 Torch/Triton 代码,已由其他任务负责迁移到 hyper_parallel/core/context_parallel/。本清理任务不负责搬运或适配,只在迁移完成后删除旧副本。

迁移任务应负责:

  • KDA/GDN 新目录代码;
  • 生产调用点切换到新路径;
  • KDA/GDN 测试和文档迁移;
  • KDA README、GDN LICENSE 的新打包路径;
  • MANIFEST.in、setup.py、Lizard 白名单的新路径;
  • 新目录 CODEOWNERS。

本清理提交必须基于迁移结果开发,不能在调用点仍引用旧路径时提前删除 platform/torch/custom_ops/。

二、当前结构及处理方式

hyper_parallel/
├── custom_ops/                                      [整体删除]
│   ├── __init__.py
│   └── experimental/
│       ├── __init__.py
│       └── experimental_ops.py
│
├── platform/
│   ├── platform.py                                  [删除 custom_ops 抽象属性]
│   │
│   ├── mindspore/
│   │   ├── platform.py                              [删除 custom_ops 属性和缓存]
│   │   └── custom_ops/                              [整体删除]
│   │       ├── __init__.py
│   │       ├── custom_ops.py
│   │       ├── custom_op_impl.py
│   │       ├── build.sh
│   │       ├── CMakeLists.txt
│   │       ├── module.cc
│   │       ├── module.h
│   │       ├── dense_lightning_indexer_*.cc
│   │       ├── sparse_lightning_indexer_*.cc
│   │       ├── lightning_indexer_v2.cc
│   │       ├── mhc_*.cc
│   │       └── sparse_flash_mla*.cc
│   │
│   └── torch/
│       ├── platform.py                              [删除 custom_ops 属性和缓存]
│       └── custom_ops/                              [迁移完成后整体删除]
│           ├── __init__.py                          [TorchCustomOps 全部未实现]
│           ├── kda/                                 [迁移后删除旧副本]
│           └── gdn/                                 [迁移后删除旧副本]
│
└── core/
    ├── context_parallel/                            [迁移任务负责,本任务不修改]
    └── shard/ops/
        ├── parallel_mhc_post.py                     [删除]
        ├── parallel_mhc_pre_sinkhorn.py             [删除]
        ├── parallel_npu_dense_lightning_indexer_grad_kl_loss.py
        │                                             [删除]
        ├── parallel_npu_dense_lightning_indexer_softmax_lse.py
        │                                             [删除]
        ├── parallel_npu_sparse_lightning_indexer_grad_kl_loss.py
        │                                             [删除]
        ├── yaml/
        │   ├── npu_mhc_post_ops.yaml                [删除]
        │   ├── npu_mhc_pre_sinkhorn_ops.yaml        [删除]
        │   ├── npu_dense_lightning_indexer_grad_kl_loss_ops.yaml
        │   │                                         [删除]
        │   ├── npu_dense_lightning_indexer_softmax_lse_ops.yaml
        │   │                                         [删除]
        │   └── npu_sparse_lightning_indexer_grad_kl_loss_ops.yaml
        │                                             [删除]
        ├── parallel_lightning_indexer.py             [保留,解除对删除文件的依赖]
        └── parallel_npu_sparse_flash_attention.py    [保留,解除对删除文件的依赖]

其中,parallel_lightning_indexer.py 和 parallel_npu_sparse_flash_attention.py 是独立的分布式算子支持,不属于要删除的 MindSpore custom_ops 实现。

它们目前仍从即将删除的 parallel_npu_dense_lightning_indexer_softmax_lse.py 导入共享 helper。清理时需要把仍被使用的 helper 收到保留的公共实现中,不能让保留代码继续依赖被删除文件。

三、测试目录清理

tests/
├── ut/
│   ├── custom_ops/                                  [整体删除]
│   │   └── experimental/
│   │       └── test_experimental_ops.py
│   │
│   ├── platform/
│   │   ├── mindspore/custom_ops/                    [整体删除]
│   │   │   └── test_custom_op_impl.py
│   │   └── torch/
│   │       └── test_kda_fla_adapter.py              [KDA 迁移任务负责]
│   │
│   └── core/shard/ops/
│       ├── test_parallel_mhc_post.py                 [删除]
│       ├── test_parallel_mhc_pre_sinkhorn.py         [删除]
│       ├── test_parallel_npu_dense_lightning_indexer_grad_kl_loss.py
│       │                                             [删除]
│       ├── test_parallel_npu_dense_lightning_indexer_softmax_lse.py
│       │                                             [删除]
│       ├── test_parallel_npu_sparse_lightning_indexer_grad_kl_loss.py
│       │                                             [删除]
│       ├── test_parallel_lightning_indexer.py        [保留并接收共享 helper UT]
│       └── test_parallel_npu_sparse_flash_attention.py
│                                                     [保留]
│
└── mindspore/st/shard/ops/cases/
    ├── case_mhc_post.py                              [删除]
    ├── case_mhc_pre_sinkhorn.py                      [删除]
    ├── case_npu_dense_lightning_indexer_grad_kl_loss.py
    │                                                 [删除]
    ├── case_npu_dense_lightning_indexer_softmax_lse.py
    │                                                 [删除]
    └── case_npu_sparse_lightning_indexer_grad_kl_loss.py
                                                      [删除]

MindSpore shard-op 框架通过扫描 case_*.py 自动注册,因此删除这五个 case 后没有额外的用例清单需要维护。

四、编译和打包清理

根构建入口

根 build.sh 删除:

CUSTOM_OPS_VALUE
--custom-ops on|off
custom_ops 组件的 run_component 调用
custom_ops 对 CANN 环境激活条件的影响

最终根构建结构保持为:

build.sh
├── 调用 multicore 组件编译
├── 调用 indexed 组件编译
├── 汇总组件 payload
└── 统一生成 wheel

推荐构建命令:

bash build.sh --multicore on --strict on

删除组件构建目录

以下内容随 hyper_parallel/platform/mindspore/custom_ops/ 整体删除,不保留独立 custom_ops 构建入口:

build.sh
CMakeLists.txt
*.cc
module.h

Docker 清理

删除 BUILD_CUSTOM_OPS_EXTENSION,涉及:

docker/
├── Dockerfile.mindspore
├── Dockerfile.torch
├── Dockerfile.hyper-parallel-npu
├── build_hyper-parallel_npu.sh
└── README.md

该变量当前只被声明和透传,没有实际接入 Docker 安装过程,删除不会减少现有有效能力。

打包配置职责

涉及:

MANIFEST.in
setup.py
.jenkins/check/config/whitelizard.txt

KDA/GDN 新路径由迁移任务维护。本清理任务只保证:

  • 不再出现 platform/torch/custom_ops 旧路径;
  • 不再出现 MindSpore custom_ops 路径;
  • 不删除迁移任务已经写入的新路径配置。

五、资料清理

需要更新:

README.en.md
RELEASE.md
RELEASE_CN.md
docs/
├── installation.md
├── faq.md
└── contributing/
    ├── dev_environment.md
    └── release.md

hyper_parallel/core/multicore/docs/build.md
docker/README.md

处理内容:

  • 删除 --custom-ops 参数;
  • 删除 MindSpore custom_ops 编译、故障排查和单独构建说明;
  • 删除当前版本“支持 custom ops 原生扩展”的能力声明;
  • 更新构建命令。

保留 hyper_parallel_v1.0.0_release_notes.md 中的 custom_ops 历史 PR 记录。它记录的是历史事实,不代表当前能力,不应篡改。

KDA README 由迁移任务移动到 core/context_parallel 对应位置,本任务只删除旧副本。

六、CI 和所有权配置

删除 MindSpore custom_ops 的静态检查豁免:

.jenkins/check/config/filter_cpplint.txt
.jenkins/check/config/filter_pylint.txt

Lizard 配置分工:

  • 迁移任务添加 KDA/GDN 新路径;
  • 清理任务删除 platform/torch/custom_ops 旧路径。

当前 CODEOWNERS 没有 custom_ops 专属规则,因此本任务无需增加已删除目录的 owner;KDA/GDN 新路径的 owner 由迁移任务负责。

.agent/ 下没有 custom_ops 专属 skill、agent 或规则,本子任务不需要修改 agent 资产。Platform 整体删除带来的规则更新归 ISSUE #377 的主任务处理。

七、目标目录状态

hyper_parallel/
├── custom_ops/                                  [不存在]
├── platform/
│   ├── mindspore/custom_ops/                    [不存在]
│   └── torch/custom_ops/                        [不存在]
└── core/
    ├── context_parallel/                        [迁移后的 KDA/GDN,由迁移任务负责]
    └── shard/ops/
        ├── legacy custom_ops parallel 文件       [不存在]
        ├── legacy custom_ops YAML                [不存在]
        ├── parallel_lightning_indexer.py         [保留]
        └── parallel_npu_sparse_flash_attention.py [保留]

tests/
├── ut/custom_ops/                               [不存在]
├── ut/platform/mindspore/custom_ops/            [不存在]
├── legacy custom_ops shard UT                   [不存在]
└── legacy custom_ops MindSpore ST               [不存在]

Platform 其他目录暂时仍存在,等待 ISSUE #377 的其他子任务继续清理。

八、验收条件

重点验证“特性消失”,而不是全仓 custom_ops 字符串清零:

  • hyper_parallel/custom_ops/、platform/mindspore/custom_ops/、platform/torch/custom_ops/ 均不存在;
  • Platform.custom_ops、MindSporeCustomOps、TorchCustomOps 不存在;
  • 10 个 hyper_parallel.custom_ops.experimental 接口不存在;
  • hyper_parallel_custom_ops_ms.so 不再编译、加载或进入安装包;
  • --custom-ops、BUILD_CUSTOM_OPS_EXTENSION 不存在;
  • legacy custom_ops 的 YAML 注册、UT、MindSpore ST 不存在;
  • KDA/GDN 调用均指向迁移后的 core/context_parallel;
  • wheel/sdist 中不存在旧目录和 MindSpore custom_ops 制品;
  • Multicore 开启和关闭两种构建均通过;
  • KDA/GDN 迁移后的测试、保留的 shard-op UT 和全量 UT 通过。

允许保留的 custom_ops 字符串包括:

  • 外部依赖 omni_training_custom_ops;
  • torch.ops.custom.*;
  • tests/ut/core/shard/custom_parallel_ops 测试夹具;
  • 普通“custom operator”说明;
  • 历史发布记录。
likedislike
Yyangzhenzhang成员
19 天前 关联了pull request:refactor(dtensor): 解耦 DTensor 核心模块与 platform 层依赖
Hhedongdong成员
19 天前 关联了pull request:refactor: remove legacy custom ops feature
hedongdong成员
18 天前 评论:

Custom_ops 清理实施进展

对应 PR:!1389

本次已完成以下清理:

  • 删除 hyper_parallel/custom_ops 实验接口和 MindSpore Custom Ops 实现、编译入口。
  • 删除 Platform 层 custom_ops 契约、缓存及 Torch 未实现占位接口。
  • 删除 ISSUE #150 对应的 LightningIndexer、SparseFlashAttention、Grad-KL 和 MHC 分布式算子实现、YAML、UT 及 MindSpore ST。
  • 删除 npu_mhc_pre_clamp_sinkhorn 的完整能力链,包括实验接口、DFunction、分布式规则、C++ 前反向实现和测试。
  • npu_mhc_pre_clamp_sinkhorn 与 npu_mhc_pre_sinkhorn 原本共用 parallel_mhc_pre_sinkhorn.py 和同一 YAML,当前两者均已下架。
  • DFunction 已改为直接继承 torch.autograd.Function,不再依赖 Platform 选择自动微分基类,同时删除 MindSpore DFunction ST 和跨框架资料承诺。
  • 根 build.sh、Docker、安装/发布资料和 CI 过滤配置中的 legacy Custom Ops 入口已清理。

当前边界:

  • KDA/GDN 旧目录仍有生产调用,等待负责迁移到 core/context_parallel 的变更合入后删除。
  • Platform 其他接口和实现不在本子任务中删除,由 ISSUE #377 后续任务继续处理。
  • 历史发布记录以及外部依赖中的 custom_ops 字样不作为清零目标。

验证结果:

  • Shard UT:1240 passed,38 subtests passed。
  • DFunction UT:16 passed。
  • DSA Context Parallel UT:27 passed。
  • DFunction Torch CPU/Gloo ST:通过。
  • YAML registry、Torch-only 导入策略、Pylint 和残留引用扫描通过。
likedislike
cuiyushicuiyushi成员
18 天前 关联了pull request:refactor(integration): remove platform backend dependency
pangxzpangxz成员
18 天前 关联了pull request:refactor: remove platform in expert parallel code
Yyangzhenzhang成员
15 天前 关联了pull request:refactor(dtensor): 抽取平台无关 DTensorBase 到 core
Yyangzhenzhang成员
15 天前 关联了pull request:test: 清理 MindSpore 用例及悬空引用
Yyangzhenzhang成员
14 天前 关联了pull request:refactor: 将 EXISTING_COMM_GROUPS 从 platform 迁移至 core.utils
Mmindspore-ci-bot成员
9 天前 关联了pull request:[mirror] mindspore-ai/hyper-parallel#911: fix: reuse loss kernel and preserve mmap index views
Mmindspore-ci-bot成员
8 天前 关联了pull request:[mirror] mindspore-ai/hyper-parallel#917: refactor: consolidate core collectives and helpers into core/utils
changzheruichangzherui成员
6 天前 关联了pull request:refactor(auto_parallel): drop MindSpore config and default to PyTorch