| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
[feature] SGLang 支持 native PD Co-authored-by: qq_40172610<chenchaofeng5@huawei.com> # message auto-generated for no-merge-commit merge: !649 merge chaofeng into master [feature] SGLang 支持 native PD Created-by: qq_40172610 Commit-by: qq_40172610 Merged-by: towncharlie Description: ## 1. 合入背景 SGLang 在 EngineServer **native CLI**( sglang.launch_server)拉起时,业务口由引擎原生进程占用,不再走 Motor InferEndpoint / _motor_dispatch 适配层。现有 Unified PD 路由仍向引擎注入 _motor_dispatch,无法与 stock SGLang PD(bootstrap_host/port/room)对齐,导致 SGLang native PD 无法正常协同。 本 PR 让 Coordinator 在识别到目标实例 engine_type=sglang 时,改为注入 SGLang 原生 PD 字段,并对无 /v1/dispatch/stop 的 native 业务口做 best-effort 跳过。 Fixes [#382](https://gitcode.com/Ascend/MindIE-Motor/issues/382) ## 2. 修改内容 1. 新增 motor/coordinator/router/sglang_native_dispatch.py:识别 SGLang 实例;由 pair_id + attempt_seq 稳定派生 bootstrap_room;注入 bootstrap_host(Prefill IP)、bootstrap_port(读 DISAGGREGATION_BOOTSTRAP_PORT)、bootstrap_room,并保证不附带 _motor_dispatch。 2. UnifiedPDRouter 构造下发请求时:若目标资源为 SGLang,走上述注入并提前返回;vLLM 等非 SGLang 路径仍保持原 _motor_dispatch 逻辑。 3. DispatchStopClient:对 SGLang native 端点跳过 dispatch stop HTTP(引擎无该接口),打 info 日志,不作为硬失败。 4. 文档补充 Native 拉起与 SGLang PD 字段约定、/v1/dispatch/stop 在 pure-native 下的行为说明。 5. 补充/调整 UT:test_sglang_native_dispatch.py、test_dispatch_stop_client.py、test_unified_pd_router.py(SGLang 注入 bootstrap、不带 _motor_dispatch)、test_main_native_launch.py(SGLang native 拉起路径)。 **部署约定**:Coordinator 与引擎侧需配置一致的 DISAGGREGATION_BOOTSTRAP_PORT。 ## 3. 资料变更 涉及。更新 docs/zh/user_guide/api/engine_server_interfaces.md:补充 Native 拉起与 SGLang PD(bootstrap_*)约定,以及 SGLang pure-native 下 dispatch stop 为 best-effort 的说明。 ## 4. 接口变更 涉及(Coordinator → SGLang 引擎请求体约定,客户面 OpenAPI 不变): - SGLang PD:下发请求携带 bootstrap_host / bootstrap_port / bootstrap_room,**不再**附带 _motor_dispatch。 - /v1/dispatch/stop:仍由 Motor InferEndpoint 提供;SGLang pure-native 业务口无此接口,Coordinator 侧跳过调用。 - vLLM PD 路径不变。 ## 5. 测试结果 | 场景 | 方法 | 结果 | | --- | --- | --- | | bootstrap_room 稳定性 / 注入字段 / 缺省端口报错 | UT:test_sglang_native_dispatch.py | 通过(提交前本地执行) | | UnifiedPD 对 SGLang 注入 bootstrap、不带 _motor_dispatch;原 vLLM concurrent 路径保持 | UT:test_unified_pd_router.py | 通过 | | SGLang 跳过 dispatch stop | UT:test_dispatch_stop_client.py | 通过 | | SGLang native CLI 拉起不走 InferEndpoint | UT:test_main_native_launch.py | 通过 | | SGLang PD 端到端(真实集群 / Ascend) | 部署 P/D + 配置 DISAGGREGATION_BOOTSTRAP_PORT 后发起推理 | 通过  | ## 6. CheckList - [x] 代码注释完备 - [x] 正确记录维测日志 - [x] 是否有UT用例 - [x] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-Motor!649 | 2 个月前 | |
add storage feature Co-authored-by: supermario_leo<leo.stack@outlook.com> # message auto-generated for no-merge-commit merge: !440 merge feature_add_ucm into master add storage feature Created-by: leo6393 Commit-by: supermario_leo Merged-by: towncharlie Description: ## **1. 合入背景** 在分布式 PD 架构上接入 **UCMConnector** 作为 KV 存储层,为 Prefill 节点提供前缀缓存(prefix cache)复用能力,降低重复前缀的 TTFT、提升整体吞吐。 由于 UCM 的 storage_backends 需要一块 **P/D 跨节点共享**的持久化存储,本分支同时补齐了通用的**引擎 Pod 存储能力**(不绑定 UCM),作为 UCM 前缀缓存池的落盘载体。 - 关联特性:UCM 前缀缓存接入(分布式 PD) - 不涉及对既有 PR 问题的修复 ## **2. 修改内容** 1. **UCM Connector 接入(motor/engine_server/core/vllm/vllm_config.py)** - MultiConnector 场景下识别 connectors[1] = UCMConnector:提前处理并强制 kv_role = kv_both(UCM 在 P/D 两端均双向 store),**不注入任何 rpc port**,避免污染 UCM 的内联 kv_connector_extra_config,处理后 return early。 - 独立 UCMConnector 在 prefill/decode 角色下 **fail loud**(该拓扑需额外 dispatch/profile 工作,当前不支持);union 角色原样透传。 - AscendStoreConnector / MooncakeStore 分支行为保持不变,报错信息修正为打印实际 connector 名。 2. **端口注入适配(motor/config/endpoint.py、motor/config/node_manager.py)** - MultiConnector 组装时,UCM store 无 lookup_rpc_port。改用 .get() 判断 connector 类型,**仅跳过 UCM**;其余 store(如 AscendStore)保持直接索引 —— 真正缺失端口时仍按原逻辑 fail fast,不被静默跳过。 3. **常量新增(motor/engine_server/constants/constants.py、motor/config/node_manager.py)** - UCM_CONNECTOR = "UCMConnector"、KV_BOTH = "kv_both"、KV_CONNECTOR_MODULE_PATH。 4. **通用引擎 Pod 存储能力(新增 examples/deployer/lib/generator/storage.py)** - motor_deploy_config.storage 为**列表**,每条目必须声明 type: - pvc —— 两种互斥模式:storage_class_name 动态创建 PVC / claim_name **挂载已有 PVC**(不生成 PVC 对象,且与 size/access_mode/storage_class_name 互斥,混填/漏填均 fail loud;claim_name 用于非 pvc 类型也会被拒绝)。 - nfs —— k8s 原生 NFS 卷直挂(server+path),无需 CSI/供给器。 - hostpath —— 节点本地目录直挂。 - dshm_size —— 独立抬高 /dev/shm emptyDir sizeLimit(容纳 UCM CacheStore),带无单位值校验。 - 挂载逻辑幂等,复用到 engine.py / infer_service.py / single_container.py 三条部署路径;卷名/动态 PVC 名按序号自动生成(mindie-motor-store-<i>)。 5. **示例与镜像** - examples/infer_engines/vllm/ucm_pd/(README + user_config.json):分布式 PD + UCM 内联配置示例。 - docker/mindie-motor-vllm-ucm/(Dockerfile + README):UCM 叠加镜像。 ## **3. 资料变更** **涉及。** - 新增 examples/infer_engines/vllm/ucm_pd/README.md:UCM 前缀缓存拓扑、存储三种类型(pvc/nfs/hostpath)字段表、动态创建 / 挂已有 PVC(claim_name)两种写法示例、dshm_size 说明及硬约束(storage_backends == mount_path)。 - 新增 docker/mindie-motor-vllm-ucm/README.md:UCM 叠加镜像构建说明。 ## **4. 接口变更** **涉及客户面可见的配置项变更(user_config.json),不涉及跨代码仓 API 变更。** - kv_connector 支持 UCMConnector,及 MultiConnector 中内嵌 UCM 作为 connectors[1];需配套 kv_connector_module_path。 - motor_deploy_config 新增 storage 列表(type: pvc/nfs/hostpath,pvc 支持 claim_name 挂已有卷)与 dshm_size。 - 均为新增可选项,对存量配置向后兼容。 ## **5. 测试结果** 新增 / 覆盖 UT: | 用例文件 | 用例数 | 覆盖场景 | |---|---|---| | tests/examples/deployer/test_storage.py | 22 | pvc/nfs/hostpath 三类挂载;pvc 动态创建 vs claim_name 挂已有;参数互斥与缺失校验;非 pvc 类型拒绝 claim_name;多条目/同 claim 多挂载点;挂载幂等;dshm_size 设置与无单位校验 | | tests/engine_server/core/test_vllm_config.py | 13 | UCM kv_role=kv_both、不注入端口;独立 UCM 在 P/D 角色 fail loud;MultiConnector 内嵌 UCM;既有 mooncake/ascend 分支回归 | | tests/node_manager/test_config.py | 33 | MultiConnector 组装跳过 UCM 端口注入;非 UCM store 缺端口仍 fail fast | - 部署方式维度:覆盖 Deployment(engine)、InferServiceSet、single-container 三条路径的存储挂载。 - 本地已执行 deployer UT 套件:54 passed;test_vllm_config.py / test_config.py 依赖运行时环境,请在集成环境执行确认。 - 建议补充:UCM 叠加镜像 + 实机分布式 PD 端到端前缀缓存命中 / TTFT 验证。 ## **6. CheckList** - [x] 代码注释完备 - [x] 正确记录维测日志 - [x] 是否有UT用例 - [x] 若涉及多线程场景,考虑了并发场景,不存在死锁问题(本次改动为配置解析 / YAML 生成,无新增多线程逻辑) See merge request: Ascend/MindIE-PyMotor!440 | 2 个月前 | |
[feature] 支持 Prefill 跨机 Pipeline Parallel 拉起 Co-authored-by: Jechin<yuzechen1@huawei.com> # message auto-generated for no-merge-commit merge: !653 merge feature/prefill-cross-node-pp into master [feature] 支持 Prefill 跨机 Pipeline Parallel 拉起 Created-by: Jechin Commit-by: Jechin Merged-by: towncharlie Description: ## **1. 合入背景** Fixes [#389](https://gitcode.com/Ascend/MindIE-Motor/issues/389) Prefill 跨机 Pipeline Parallel(如 TP=16、PP=2、nnodes=2)场景下,控制面仍按“单机并行积”计算 local_world_size(pcp×tp×pp),导致本机设备数校验失败、Endpoint 为空;同时 master_addr=placeholder / 静态 node_rank 会盖住运行时注入,分布式初始化无法连通。Assembler 在无 Endpoint 时还会误报 start 成功。本 PR 补齐跨机 PP 与跨机 PCP 共用的 nnodes 路径,并让 Deployer 从 vLLM 脚本正确生成 PP/nnodes 配置;并行度读取统一为 CLI 整数优先、缺省回退 kv_connector_extra_config(含 pp_size)。 ## **2. 修改内容** 1. **NodeManager 配置**(motor/config/node_manager.py) - 跨机时按 (pcp×tp×pp)//nnodes 折算本机 local_world_size(覆盖 PP/PCP) - pcp×pp 不能被 nnodes 整除时直接报错,避免错误拓扑静默通过 2. **EngineServer VLLMConfig**(motor/engine_server/core/vllm/vllm_config.py) - 跨机场景强制覆盖 master_addr / node_rank / headless,不再被 placeholder 挡住 - Mooncake kv extra 合并并行度时保留用户字段(如 pp_layer_partition) 3. **Controller InstanceAssembler**(motor/controller/core/instance_assembler.py) - 所有 NodeManager 均无 Endpoint 时,_send_start_command 返回失败并打 ERROR,禁止假成功 4. **Deployer 转换**(examples/deployer/config_tool/vllm_to_motor.py) - 保留并正确写出 pipeline_parallel_size,按 tp×pp 推导 Pod / nnodes - 跨机时写入 nnodes / master-port;不写 master-addr / node-rank(运行时注入) - **dp/tp/pp 统一读取优先级**:命令行给了正整数用 CLI,否则回退 kv_connector_extra_config 的 dp_size / tp_size / pp_size(kv 中的 size 在写出前剥离) - 去除硬件侧强制 remap tp/dp;infer_*_motor_deploy_config 回传 nnodes,与跨机 engine 注入共用一次 _infer_pod_layout - 抽取 hybrid 默认 deploy / 权重挂载路径 / preset+cards 等重复逻辑,降低漂移风险 5. **UT** - 覆盖 PP 折算、不可整除、placeholder 覆盖、kv 字段保留、空 Endpoint start 失败 - Deployer:CLI>kv、仅 kv 回退(含 pp_size)、跨机 nnodes/master-port、跳过脚本透传多机键等 **进程视图** mermaid flowchart LR subgraph deploy [Deploy] script["vLLM serve 脚本\nCLI 与 kv extra"] conv["vllm_to_motor\nCLI大于kv"] uc["user_config\nPP nnodes master-port"] end subgraph control [Control Plane] nm0["NodeManager node_rank=0"] nm1["NodeManager node_rank=1"] asm["InstanceAssembler"] end subgraph engine [Engine] es0["EngineServer PP stage0"] es1["EngineServer PP stage1 headless"] end script --> conv --> uc uc --> nm0 uc --> nm1 nm0 -->|"Register local_world_size"| asm nm1 -->|"Register local_world_size"| asm asm -->|"StartCmd master_dp_ip/node_rank"| nm0 asm -->|"StartCmd"| nm1 nm0 --> es0 nm1 --> es1 es0 <-->|"master-port rendezvous"| es1 ## **3. 资料变更** 不涉及仓库内用户文档更新(Deployer README / 跨机说明未改)。 ## **4. 接口变更** 涉及配置约定(客户面可见): - Prefill 跨机 PP 需配置 pipeline_parallel_size、nnodes、master-port;**不要**配置 master-addr / node-rank - Deployer 从 vLLM 脚本转换时会自动生成上述项;pipeline_parallel_size 不再被强制改写为 1 - Deployer 并行度语义:--data/tensor/pipeline-parallel-size 正整数优先于 kv extra 的 dp_size/tp_size/pp_size;二者皆无时 dp/tp 走手动占位提示,pp 缺省为 1 - 无新增/变更 HTTP API ## **5. 测试结果** (自行补充) ## **6. CheckList** [x] 代码注释完备 [x] 正确记录维测日志 [x] 是否有UT用例 [x] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-Motor!653 | 1 个月前 | |
【重构】EngineServer重构&EngineServer对接SGLang Co-authored-by: ganglv<lvgang1@huawei.com> # message auto-generated for no-merge-commit merge: !36 merge pymotor_master_refactor_sglang into master 【重构】EngineServer重构&EngineServer对接SGLang Created-by: ganglv Commit-by: ganglv Merged-by: towncharlie Description: ## **1. 合入背景** > 请描述为什么要做这个PR内的改动。\ > 如涉及,请关联前序PR或同特性/需求下的其他PR。\ > 如果是修复之前PR引入的问题,请关联引入问题的PR。\ > 请通过#ISSUE ID关联issue。\ > 注意: Fixes #ISSUE ID会自动关闭issue,如问题部分解决请不要使用Fixes,可以用Fix part of #ISSUE ID替代. ## **2. 修改内容** > 请<ins>**描述修改内容的具体实现**</ins>,涉及哪些组件之间进行交互,可以用1、2、3、...进行罗列。 > 如果是需求或者重构类的PR,需要<ins>**补充详细设计文档**</ins>(说明上下游组件关系、时序图、类图、DFX能力等内容)。 ## **3. 资料变更** > 请确认<ins>**是否涉及资料变更**</ins>。\ > 如涉及,需要在PR中体现,并简要说明修改内容。\ > 如不涉及,需填写“不涉及”。 ## **4. 接口变更** > 请确认<ins>**是否涉及跨代码仓或者客户面可见的接口变更**</ins>。\ > 如涉及,需详细说明接口以及对应的变更内容,同时需要在资料中体现。\ > 如不涉及,需填写“不涉及”。 ## **5. 测试结果** > 需体现<ins>**测试场景,测试方法以及测试结果**</ins>。\ > 测试用例设计时需考虑硬件、部署方式、功能、性能、精度、显存等维度。 ## **6. CheckList** > PR提交人对以下CheckList自检项进行全量自检,自检通过或不涉及,均修改 [ ] 为 [x] [ ] 代码注释完备 [ ] 正确记录维测日志 [ ] 是否有UT用例 [ ] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-PyMotor!36 | 6 个月前 | |
[bugfix]去除host绑定限制 & 删除sp-block默认值 Co-authored-by: ganglv<lvgang1@huawei.com> # message auto-generated for no-merge-commit merge: !506 merge bind_host into master [bugfix]去除host绑定限制 & 删除sp-block默认值 Created-by: ganglv Commit-by: ganglv Merged-by: towncharlie Description: ## **1. 合入背景** > 请描述为什么要做这个PR内的改动。\ > 如涉及,请关联前序PR或同特性/需求下的其他PR。\ > 如果是修复之前PR引入的问题,请关联引入问题的PR。\ > 请通过#ISSUE ID关联issue。\ > 注意: Fixes #ISSUE ID会自动关闭issue,如问题部分解决请不要使用Fixes,可以用Fix part of #ISSUE ID替代. 1. docker only场景,不能通过其他宿主机访问推理服务,motor需要绑定0.0.0.0 2. 去除sp-block默认值,当前如果yaml中sp-block有值,不再自动生成 ## **2. 修改内容** > 请<ins>**描述修改内容的具体实现**</ins>,涉及哪些组件之间进行交互,可以用1、2、3、...进行罗列。 > 如果是需求或者重构类的PR,需要<ins>**补充详细设计文档**</ins>(说明上下游组件关系、时序图、类图、DFX能力等内容)。 去掉host绑定的校验 ## **3. 资料变更** > 请确认<ins>**是否涉及资料变更**</ins>。\ > 如涉及,需要在PR中体现,并简要说明修改内容。\ > 如不涉及,需填写“不涉及”。 不涉及 ## **4. 接口变更** > 请确认<ins>**是否涉及跨代码仓或者客户面可见的接口变更**</ins>。\ > 如涉及,需详细说明接口以及对应的变更内容,同时需要在资料中体现。\ > 如不涉及,需填写“不涉及”。 不涉及 ## **5. 测试结果** > 需体现<ins>**测试场景,测试方法以及测试结果**</ins>。\ > 测试用例设计时需考虑硬件、部署方式、功能、性能、精度、显存等维度。  ## **6. CheckList** > PR提交人对以下CheckList自检项进行全量自检,自检通过或不涉及,均修改 [ ] 为 [x] [ ] 代码注释完备 [ ] 正确记录维测日志 [ ] 是否有UT用例 [ ] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-PyMotor!506 | 2 个月前 | |
update license Co-authored-by: y1lou<louyi6@huawei.com> # message auto-generated for no-merge-commit merge: !185 merge update_license into master update license Created-by: y1lou Commit-by: y1lou Merged-by: ascend-robot Description: ## **1. 合入背景** > 请描述为什么要做这个PR内的改动。\ > 如涉及,请关联前序PR或同特性/需求下的其他PR。\ > 如果是修复之前PR引入的问题,请关联引入问题的PR。\ > 请通过#ISSUE ID关联issue。\ > 注意: Fixes #ISSUE ID会自动关闭issue,如问题部分解决请不要使用Fixes,可以用Fix part of #ISSUE ID替代. ## **2. 修改内容** > 请<ins>**描述修改内容的具体实现**</ins>,涉及哪些组件之间进行交互,可以用1、2、3、...进行罗列。 > 如果是需求或者重构类的PR,需要<ins>**补充详细设计文档**</ins>(说明上下游组件关系、时序图、类图、DFX能力等内容)。 ## **3. 资料变更** > 请确认<ins>**是否涉及资料变更**</ins>。\ > 如涉及,需要在PR中体现,并简要说明修改内容。\ > 如不涉及,需填写“不涉及”。 ## **4. 接口变更** > 请确认<ins>**是否涉及跨代码仓或者客户面可见的接口变更**</ins>。\ > 如涉及,需详细说明接口以及对应的变更内容,同时需要在资料中体现。\ > 如不涉及,需填写“不涉及”。 ## **5. 测试结果** > 需体现<ins>**测试场景,测试方法以及测试结果**</ins>。\ > 测试用例设计时需考虑硬件、部署方式、功能、性能、精度、显存等维度。 ## **6. CheckList** > PR提交人对以下CheckList自检项进行全量自检,自检通过或不涉及,均修改 [ ] 为 [x] [ ] 代码注释完备 [ ] 正确记录维测日志 [ ] 是否有UT用例 [ ] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-pyMotor!185 | 8 个月前 |