已合并
Feat: HyperParallel支持LLM推理Generate流程 #2101 #886
Feat: HyperParallel支持LLM推理Generate流程 #2101 #886
已合并
Moy创建于 6月22日
Moy
Moy
6月22日

What type of PR is this?
/kind feature


What does this PR do / why do we need it:

任务:【开源实习】HyperParallel支持LLM推理Generate流程 #2101
地址:https://gitcode.com/mindspore/community/issues/2101

本 PR 新增 LLM 推理 Generate 基础流程,提供模型无关的 prefill + decode 生成闭环,并提供采样、KV cache、batch 生成、分布式推理边界、真实模型对齐和性能基线验证。

主要内容:

  • 新增 hyper_parallel.infer.generate,支持 causal LM 的 prefill + autoregressive decode。
  • 新增 GenerationConfig、GenerateMixin、采样 helper、mask/position helper。
  • 支持 greedy、top-k、top-p、repetition penalty、logits processor、stopping criteria。
  • 支持 batch left-padding、attention mask、per-item EOS、padded output。
  • 新增 KVCache,支持本地 KV cache update、merge、clear、cache decode 与 no-cache fallback。
  • 支持 HuggingFace 新版 opaque Cache object 原样传递。
  • 新增 ContextParallelKVCache,支持 CP local KV shard 元数据、local cache 校验与 generate 主循环接入。
  • 支持 context_parallel_cache=True 的 CP cache generate opt-in 路径,包含 full cache 切分、local cache decode 更新、context_logits_rank 采样 logits handoff。
  • 支持 contiguous prefix cache reuse,返回结果只包含当前 input_ids 与新生成 tokens,不包含复用的 prefix tokens。
  • 新增 TP vocab-sharded logits gather,支持采样前聚合 tensor-parallel logits。
  • 新增 Qwen3.5、Qwen3.5-MoE 等仓库项目模型 generate API 兼容性测试。
  • 新增 HuggingFace 正式模型对齐脚本,验证真实文本模型的 HF model.generate、HyperParallel cache generate、HyperParallel no-cache generate token id 一致。
  • 新增轻量 benchmark 脚本,输出 prefill latency 与 decode throughput 基线。
  • 新增 CP generate 端到端检查脚本,验证 CP-aware attention 模型下 CP cache generate 与单进程 baseline 输出一致。

能力边界:

  • 普通 generate 支持 cache 与 no-cache 两条路径;当模型不返回 past_key_values 时自动走 no-cache fallback。
  • 当前 Qwen 仓库项目模型测试覆盖 no-cache fallback 兼容性;真实 KV cache generate 由 cache-capable 测试模型、benchmark、HuggingFace 正式模型对齐脚本覆盖。
  • CP generate 当前支持 contiguous sequence shard、append-only cached decode、contiguous prefix cache reuse;模型需要按 sequence_shard_info/global_seq_len 契约接收 local past_key_values 并返回更新后的 local cache 元数据。
  • CP generate 暂不支持 packed/non-contiguous cache、prefix stitching、多段 cache reuse。

Which issue(s) this PR fixes:


Test Plan and Test result:What scenarios were tested, and what were the verification results(Function, performance, reliability, etc.):

测试环境:

  • Ascend 910B
  • Python 3.10.20
  • torch: 2.6.0+cpu
  • torch_npu: 2.6.0.post3
  • mindspore: 2.8.0
  • CANN: 8.5.1

功能测试:

python -m pytest tests/torch/generate -q
44 passed, 18 warnings in 4.92s

覆盖范围:

  • GenerationConfig 参数校验。
  • greedy、top-k、top-p、repetition penalty。
  • top-k + top-p 组合采样顺序。
  • logits processor 与 stopping criteria。
  • cache decode 与 no-cache fallback。
  • HuggingFace opaque Cache object 传递。
  • cache/no-cache 生成 token id 一致性。
  • cache/no-cache decode logits cosine similarity。
  • batch left-padding、attention mask、per-item EOS、padded output。
  • 模型 forward 不支持部分 kwargs 时的参数过滤。
  • 模型内部真实 TypeError 不被 fallback 吞掉。
  • TP logits gather 与不均匀 vocab shard 显式报错。
  • CP final-token logits handoff 与 invalid context_logits_rank 显式报错。
  • CP KV cache full/local update、merge、clear、metadata 校验。
  • contiguous prefix cache reuse、prefix cache batch size 校验、模型不返回 cache 时显式报错。
  • Qwen3.5、Qwen3.5-MoE 等仓库项目模型 generate API 兼容性。

HF真实模型对齐验证:

模型:Qwen3-4B-Instruct-2507

max_new_tokens: 128
logits_compare_steps: 128
hf_vs_hyper_cache_ids_match: true
hyper_cache_vs_no_cache_ids_match: true
logits_cosine_min: 0.9996622800827026
logits_cosine_threshold: 0.999
logits_cosine_pass: true

f1.png
f2.png
(截图说明:第一张图为 hf_alignment_check.py 运行日志,第二张图为输出 JSON关键字段摘要。)

验证方式:

  • 使用 greedy decoding(do_sample=False)。
  • 对比 HuggingFace model.generate 与 HyperParallel generate(use_cache=True) 的完整 token id 序列,结果完全一致。
  • 对比 HyperParallel generate(use_cache=True) 与 generate(use_cache=False) 的完整 token id 序列,结果完全一致。
  • 对比真实模型 cache/no-cache decode logits cosine similarity,最小值高于阈值 0.999。

长序列压力对齐验证:

max_new_tokens: 1024
generated_new_tokens: 1024
logits_compare_steps: 1024
hf_vs_hyper_cache_ids_match: true
hyper_cache_vs_no_cache_ids_match: true
logits_cosine_min: 0.9999998211860657
logits_cosine_threshold: 0.999
logits_cosine_pass: true

单卡 NPU benchmark:

脚本:examples/generate/benchmark_generate.py

prefill_tokens_per_second: 1766522.11346752
with_cache_tokens_per_second: 40986.291279002704
no_cache_tokens_per_second: 39857.82083651311

with_cache output_shape: [4, 96]
no_cache output_shape: [4, 96]
with_cache.first_call: [32, 0]
with_cache.last_call: [1, 94]
no_cache.first_call: [32, 0]
no_cache.last_call: [95, 0]

双卡 distributed boundary 验证:

tp_shape: (2, 8)
tp_argmax: [6, 6]

cp_shape: (2, 4)
cp_argmax: [1, 2]

CP generate 端到端验证:

脚本:examples/generate/cp_generate_check.py

world_size: 2
max_new_tokens: 4
all_ranks_match_baseline: true
prefix_all_ranks_match_baseline: true

该验证使用 CP-aware tiny attention model,真实走 generate() 的 CP cache 路径,包括:

  • prefill 后 CP local KV shard;
  • cached decode local cache 更新;
  • context_logits_rank logits handoff;
  • contiguous prefix cache reuse;
  • CP 输出与单进程 baseline exact match。

补充验证(见下图):
在同一 Qwen3-4B-Instruct-2507 模型上,通过 Python API 直接调用hyper_parallel.infer.generate,输入 用一段话简单介绍Mindspore。,输出 decoded text。该截图展示真实提示词生成效果,正式对齐验收以上方 hf_alignment_check.py 的 token id 一致性和 logits cosine 结果为准。

f3.png

Self-checklist:(请自检,在[ ]内打上x,我们将检视你的完成情况,否则会导致pr无法合入)

likedislike
Pull Request已成功合入, 合并人@MindSpore-Bot
(感谢 Moy 的贡献)
MoyMoy
6月22日 创建了 pull request,commit 6561bba4
MoyMoy
6月22日 关联了issue:【开源实习】HyperParallel支持LLM推理Generate流程
MindSpore-BotMindSpore-Bot成员
6月22日 添加了label:mindspore-cla/yes
司小南(机器人)
司小南(机器人)成员
6月22日 评论:

@xuxinkun_2026, 当前/check-pr未通过,原因如下:

以下Pull Request描述检查项未通过:
选项未通过检查: 设计:PR对应的方案是否已经经过Maintainer评审,方案检视意见是否均已答复并完成方案修改
选项未通过检查: 测试:PR中的代码是否已有UT/ST测试用例进行充分的覆盖,新增测试用例是否随本PR一并上库或已经上库
选项未通过检查: 验证:PR描述信息中是否已包含对该PR对应的Feature、Refactor、Bugfix的预期目标达成情况的详细验证结果描述

请修改好上述检查错误后,重新使用/check-pr触发检查。

likedislike
MindSpore-BotMindSpore-Bot成员
6月22日 添加了label:no-pass-all-review
此处折叠了191条消息 查看更多
Yyao_yf成员
6月29日 通过审查
MindSpore-Bot
MindSpore-Bot成员
6月29日 评论:

Notice

The PR needs 3 assignees to review. 1 does not review. if all are passed, please comment /check-pr to try merge the PR. 😄

likedislike
Yyangzhenzhang成员
6月29日 通过审查
MindSpore-BotMindSpore-Bot成员
6月29日 删除了label:no-pass-all-review
MindSpore-BotMindSpore-Bot成员
6月29日 合入了pull request,合并节点 SHA:a850c639395d2cf576168107f032321df8fb03fb