已关闭
【RFC】HyperParallel支持qwen3_vl #1
yide12创建于  2025年12月27日关闭于  5月19日
yide12
yide12成员
2025年12月27日 创建
name about labels
RFC Use this template for requirement to be discussed kind/feature or kind/enhancement
Requirement Use this template for Confirmed requirements kind/feature or kind/enhancement

Backgroud(背景信息)

  • Describe/Explain the status of the problem you wish to solve.
  • Attach relevant issues if there is any.

HyperParallel是新一代并行加速库,将支持大模型库transformers的主流模型。

Origin(信息来源)

  • Explain which department/team made this request so that its priority can be given.

Benefit / Necessity (价值/作用)

  • Describe/Explain the key value by fulfilling the request.

支持transformers主流模型,提高大模型领域市场竞争力。

Design(设计方案)

  • Describe/Explain the general idea of the design. Pseudo-code is allowed

HpyerParallel支持transformer主流模型:

  1. 支持相关分布式算子。
  • 统计缺失算子。
  • 补齐缺失算子。
  1. 支持相关并行功能。
  • shard
  • hsdp
  • xxx

支持相关分布式算子

原理:算子分发机制

HyperParallel,对标torch实现了一套 _op_dispatch 机制,用于将DTensor转换成普通tensor+通信操作。

其中,DTensor通过op_dispatch拦截所有算子调用,流程如下:

  1. Op Dispatcher: 拦截算子(如 add(a, b))。
  2. Unwrap: 将输入的 DTensor 解包为本地tensor和切分元信息。
  3. 切分推导:
    • 根据输入的layout和算子规则(定义在 tensor/ops/ 目录下yaml),推导输出应该是什么layout。
    • 例如:两个输入都是行切,输出layout也是行切。
    • 如果输入不兼容,一个行切、一个列切,就会决定插入通信操作。
  4. Redistribute (如果需要):
    • 如果推导出的策略需要通信,则在本地 Tensor 上执行相应的通信操作。
  5. Local Execution: 在本地 Tensor 上执行原始算子。
  6. Wrap: 将结果 Local Tensor 重新包装为 DTensor 返回给用户。
算子分发流程图:

算子分发流程图

实现:transformers接入HyperParallel

简单代码示例:通过将输入input转换为hyperParallel的DTensor,从而接入hyperParallel框架。

# 从transformer导入模型并初始化
from transformers import AutoProcessor, AutoModelForVision2Seq
model_id = "/data3/w30032396/run_net/qwen3_vl/model/Qwen3-VL-4B-Instruct"
processor = AutoProcessor.from_pretrained(model_id, local_files_only=True, device_map="auto")
model = AutoModelForVision2Seq.from_pretrained(model_id, local_files_only=True, device_map="auto", attn_implementation="flash_attention_2")
# 应用hyperParallel的切分策略
# from hyper_parallel import shard, hsdp
# shard(model, shard_strategy)
# model = hsdp(model)

# 输入转为hyperParallel的DTensor,用于接入hyperParallel的分布式算子流程
# 这里仅为统计缺失分布式算子,单卡运行,不进行dtp切分。
from hyper_parallel import DTensor, Layout
layout = Layout((4, 2), ("dp", "tp"))
input_layout = layout("None", "None")
inputs.data["input_ids"] = DTensor.from_local(inputs.data["input_ids"], input_layout)
inputs.data["attention_mask"] = DTensor.from_local(inputs.data["attention_mask"], input_layout)
inputs.data["pixel_values"] = DTensor.from_local(inputs.data["pixel_values"], input_layout)
inputs.data["image_grid_thw"] = DTensor.from_local(inputs.data["image_grid_thw"], input_layout)

# 前向传播
outputs = model.generate(**inputs, max_new_tokens=40)
统计网络缺失算子结果

qwen3_vl:缺失算子59个。qwen3_omni:缺失算子63个。两个网络缺失算子绝大部分是重复的,去重后,缺失算子总数为75个。

详细统计信息见表格:https://onebox.huawei.com/v/8cdaebd3528e36bafcda5097dd6e8729?type=0

likedislike
yide12yide12成员
2025年12月27日 创建了RFC
yide12yide12成员
2025年12月27日 修改了描述
yide12yide12成员
2025年12月27日 修改了描述
yide12yide12成员
2025年12月27日 修改了描述
yide12yide12成员
2025年12月27日 修改了描述
yide12yide12成员
2025年12月27日 修改了描述
yide12yide12成员
2025年12月27日 修改了描述
yide12yide12成员
1月26日 修改标题为 “【RFC】HyperParallel支持qwen3_vl”,原标题为“HyperParallel支持qwen3_vl”
jiangna1111jiangna1111成员
4月9日 关联了pull request:feat(dcp): minimal final-plan cache fast-path and planner-owned save knobs
lishannilishanni
4月9日 关联了pull request:fix: improve tensor parallel plan validation and observability
yide12yide12成员
5月19日 issue状态由 TODO 改变为 DONE
yide12yide12成员
5月19日 关闭了 issue