已开启
[Bug]: Importing non-Omni high-performance modules requires optional Omni ops #369
cuiyushi创建于  9月7日
cuiyushi
cuiyushi成员
9月7日 创建

Checklist

🐛 Describe the bug

hyper_parallel.components.modules.__init__ eagerly imports every high-performance module.
Some modules, such as DSA/MHC, use the optional omni_training_custom_ops package. As a
result, importing an unrelated module such as RMSNorm can fail in an environment without
Omni custom ops, even though RMSNorm itself only uses torch_npu.npu_rms_norm.

Minimal reproduction in an environment without omni_training_custom_ops:

from hyper_parallel.components.modules import RMSNorm

The package import should not load DSA/MHC or require Omni for this usage.

Expected behavior

Public high-performance modules should be resolved lazily. Importing and using modules that
do not depend on Omni, including Qwen3-MoE's default RMSNorm, GQAAttention, and
GroupedExperts replacements, should work without installing omni_training_custom_ops.
The dependency should be checked only when an Omni-backed function or module is selected.

Additional context

The functional package already resolves public functions lazily. The modules package needs
the same import boundary. A fresh-process validation that blocks omni_training_custom_ops
passes for top-level imports, functional/modules imports, Qwen3-MoE registration and RMSNorm
replacement after the fix.

Environment info

  • Python 3.11
  • PyTorch/torch_npu environment: veomni_cys
  • Transformers 5.5.3 for the Qwen3-VL/Qwen3-MoE validation
  • Ascend NPU, 8-card DP/FSDP validation
likedislike
cuiyushicuiyushi成员
9月7日 添加了label:bug
庄昭雄
庄昭雄
2 天前 评论:

这条在当前的 master 上已经是修好的状态了,记一下位置,方便关闭。

hyper_parallel/components/modules/__init__.py 现在是惰性解析的:

  • docstring 直接写了原因:「Exports are resolved lazily so modules backed only by torch-npu do not import unrelated optional Omni custom operators.」
  • :28-43 是 _EXPORT_TO_MODULE,把每个公开名映射到它所属的子模块("RMSNorm": "rms_norm"、"MhcPreModule": "mhc" 等);
  • :47-55 用 __getattr__(PEP 562)按需 importlib.import_module(f"hyper_parallel.components.modules.{submodule}"),只加载被访问的那个子模块,并缓存进 globals()。

所以 from hyper_parallel.components.modules import RMSNorm 不会再牵出 omni_training_custom_ops,issue 里描述的「eagerly imports every high-performance module」已经不成立。

我没有实测这一步:当前环境里 torch 本身没有安装(hyper_parallel/__init__.py:39 会连带 import 它),所以跑不了那句 import。判断依据是上面这段代码本身。

likedislike