已合并
Feat: HyperParallel支持LLM推理Generate流程 #2101 #886
Moy创建于 6月22日
Feat: HyperParallel支持LLM推理Generate流程 #2101 #886
已合并
Pull Request已成功合入, 合并人@MindSpore-Bot
(感谢 Moy 的贡献)6月22日 关联了issue:【开源实习】HyperParallel支持LLM推理Generate流程
6月22日 添加了label:mindspore-cla/yes
司小南(机器人)
6月22日 评论:
6月22日 评论:
@xuxinkun_2026, 当前/check-pr未通过,原因如下:
以下Pull Request描述检查项未通过:
选项未通过检查: 设计:PR对应的方案是否已经经过Maintainer评审,方案检视意见是否均已答复并完成方案修改
选项未通过检查: 测试:PR中的代码是否已有UT/ST测试用例进行充分的覆盖,新增测试用例是否随本PR一并上库或已经上库
选项未通过检查: 验证:PR描述信息中是否已包含对该PR对应的Feature、Refactor、Bugfix的预期目标达成情况的详细验证结果描述
请修改好上述检查错误后,重新使用/check-pr触发检查。


6月22日 添加了label:no-pass-all-review
此处折叠了191条消息 查看更多
MindSpore-Bot
6月29日 评论:
6月29日 评论:
6月29日 通过审查
6月29日 删除了label:no-pass-all-review
6月29日 合入了pull request,合并节点 SHA:a850c639395d2cf576168107f032321df8fb03fb
What type of PR is this?
/kind feature
What does this PR do / why do we need it:
任务:【开源实习】HyperParallel支持LLM推理Generate流程 #2101
地址:https://gitcode.com/mindspore/community/issues/2101
本 PR 新增 LLM 推理 Generate 基础流程,提供模型无关的 prefill + decode 生成闭环,并提供采样、KV cache、batch 生成、分布式推理边界、真实模型对齐和性能基线验证。
主要内容:
hyper_parallel.infer.generate,支持 causal LM 的 prefill + autoregressive decode。GenerationConfig、GenerateMixin、采样 helper、mask/position helper。KVCache,支持本地 KV cache update、merge、clear、cache decode 与 no-cache fallback。ContextParallelKVCache,支持 CP local KV shard 元数据、local cache 校验与 generate 主循环接入。context_parallel_cache=True的 CP cache generate opt-in 路径,包含 full cache 切分、local cache decode 更新、context_logits_rank采样 logits handoff。input_ids与新生成 tokens,不包含复用的 prefix tokens。model.generate、HyperParallel cache generate、HyperParallel no-cache generate token id 一致。能力边界:
past_key_values时自动走 no-cache fallback。sequence_shard_info/global_seq_len契约接收 localpast_key_values并返回更新后的 local cache 元数据。Which issue(s) this PR fixes:
Test Plan and Test result:What scenarios were tested, and what were the verification results(Function, performance, reliability, etc.):
测试环境:
功能测试:
覆盖范围:
GenerationConfig参数校验。TypeError不被 fallback 吞掉。context_logits_rank显式报错。HF真实模型对齐验证:
模型:
Qwen3-4B-Instruct-2507(截图说明:第一张图为
hf_alignment_check.py运行日志,第二张图为输出 JSON关键字段摘要。)验证方式:
do_sample=False)。model.generate与 HyperParallelgenerate(use_cache=True)的完整 token id 序列,结果完全一致。generate(use_cache=True)与generate(use_cache=False)的完整 token id 序列,结果完全一致。0.999。长序列压力对齐验证:
单卡 NPU benchmark:
脚本:
examples/generate/benchmark_generate.py双卡 distributed boundary 验证:
CP generate 端到端验证:
脚本:
examples/generate/cp_generate_check.py该验证使用 CP-aware tiny attention model,真实走
generate()的 CP cache 路径,包括:context_logits_ranklogits handoff;补充验证(见下图):
在同一 Qwen3-4B-Instruct-2507 模型上,通过 Python API 直接调用
hyper_parallel.infer.generate,输入用一段话简单介绍Mindspore。,输出 decoded text。该截图展示真实提示词生成效果,正式对齐验收以上方hf_alignment_check.py的 token id 一致性和 logits cosine 结果为准。Self-checklist:(请自检,在[ ]内打上x,我们将检视你的完成情况,否则会导致pr无法合入)