| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
coordinator多进程版本合入主干 Co-authored-by: ganglv<lvgang1@huawei.com> Co-authored-by: tobking<wangjun292@huawei.com> Co-authored-by: j00813896<jiangwentao7@huawei.com> # message auto-generated for no-merge-commit merge: !251 merge br_home_base_multiprocess into master coordinator多进程版本合入主干 Created-by: tobking Commit-by: tobking;j00813896;ganglv Merged-by: ascend-robot Description: ## **1. 合入背景** > 请描述为什么要做这个PR内的改动。\ > 如涉及,请关联前序PR或同特性/需求下的其他PR。\ > 如果是修复之前PR引入的问题,请关联引入问题的PR。\ > 请通过#ISSUE ID关联issue。\ > 注意: Fixes #ISSUE ID会自动关闭issue,如问题部分解决请不要使用Fixes,可以用Fix part of #ISSUE ID替代. coordinator优化为多进程版本,提升高QPS场景下推理性能 [#158](https://gitcode.com/Ascend/MindIE-pyMotor-private/issues/158) ## **2. 修改内容** > 请<ins>**描述修改内容的具体实现**</ins>,涉及哪些组件之间进行交互,可以用1、2、3、...进行罗列。 > 如果是需求或者重构类的PR,需要<ins>**补充详细设计文档**</ins>(说明上下游组件关系、时序图、类图、DFX能力等内容)。 1.coordinator拆分为mgmt,scheduler,infer 三种独立进程 2.coordinator主进程为deamon进程,其负责拉起上述三种子进程 ## **3. 资料变更** > 请确认<ins>**是否涉及资料变更**</ins>。\ > 如涉及,需要在PR中体现,并简要说明修改内容。\ > 如不涉及,需填写“不涉及”。 涉及,userconfig新增多进程相关配置项 ## **4. 接口变更** > 请确认<ins>**是否涉及跨代码仓或者客户面可见的接口变更**</ins>。\ > 如涉及,需详细说明接口以及对应的变更内容,同时需要在资料中体现。\ > 如不涉及,需填写“不涉及”。 不涉及 ## **5. 测试结果** > 需体现<ins>**测试场景,测试方法以及测试结果**</ins>。\ > 测试用例设计时需考虑硬件、部署方式、功能、性能、精度、显存等维度。 ## **6. CheckList** > PR提交人对以下CheckList自检项进行全量自检,自检通过或不涉及,均修改 [ ] 为 [x] [x] 代码注释完备 [x] 正确记录维测日志 [x] 是否有UT用例 [x] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-pyMotor-private!251 | 7 个月前 | |
[fix] 修复 SGLang 单端口多 DP 时 KV 亲和选不到已缓存实例的问题 Co-authored-by: qq_40172610<chenchaofeng5@huawei.com> # message auto-generated for no-merge-commit merge: !1036 merge fix/sglang-l1-kv-affinity into master [fix] 修复 SGLang 单端口多 DP 时 KV 亲和选不到已缓存实例的问题 Created-by: qq_40172610 Commit-by: qq_40172610 Merged-by: tobking Description: ## **1. 合入背景** Motor KV 亲和原先按 vLLM 的「一个 HTTP endpoint 对应一个 DP」来注册和打分。SGLang 官方 launch_server --dp-size 是 **一个 Tokenizer HTTP + 进程内 DataParallelController**,每个 DP 单独发 ZMQ KV 事件(5557+rank)。 在 SGLang 2P1D、dp-size=16、enable_multi_endpoints=false 下会出现: - Coordinator 只给 DP0 订 ZMQ,其它 DP 的 L1 事件进不了 kv-conductor。 - 查询命中后只读 endpoint.id=0,DP>0 上的前缀对调度不可见,热请求回退 load_balance。 - Coordinator 侧没有 vLLM 时 DSV4 tokenizer 加载失败,token_ids 为空,query 永远 miss。 - 引擎 ZMQ 被同时登记成 cpu/disk,conductor 当成 pool 流,L1 的 GPU BlockStored 进不了 HBM 树。 本 PR 只修 **L1(HBM)亲和**,让 SGLang 单 HTTP 多 DP 能按实例选到有缓存的 P。不包含 L2/HiCache。 Fixes:[#672](https://gitcode.com/Ascend/MindIE-Motor/issues/672) ## **2. 修改内容** 交互链不变:引擎 ZMQ 发 BlockStored → kv-conductor 建 HBM 树 → Coordinator tokenize 后 /query → kv_cache_affinity 给 P 实例打分 → 转发到该实例唯一 HTTP。本 PR 补的是 SGLang 拓扑下「订全、查对、选到实例」三处。 1. **conductor_api_client:按 dp_size 展开 ZMQ 订阅** 新增 _hbm_event_targets。实例只有 1 个 HTTP、但 parallel_config.dp_size>1 时,用同一 IP 为 rank 0..dp_size-1 注册 5557+rank / 6667+rank。vLLM 多 HTTP(endpoint 数 ≥ dp_size)仍走原来的 1:1。 2. **medium_endpoints 默认只报 npu** 不再把引擎口抄到 cpu/disk。未显式配置 pool 口时,subscriber 是 EventSource::Engine,L1 BlockStored(medium=GPU) 能进 HBM 树。YuanRong 等显式配了 cpu/disk 的路径不变。 3. **kv_cache_affinity:单 HTTP 时折叠各 DP 命中** 新增 _matched_raw_for_endpoint。get_all_endpoints() 长度为 1 时,对 conductor 返回的所有 DP 取加权最长命中,再给这个 HTTP 打分。多 endpoint 实例仍按 endpoint.id 对齐。 4. **Tokenizer:SGLang + DSV4 能算出和引擎一致的 token** DSV4 不再要求 engine_type=="vllm" 才加载;无 DeepseekV4Tokenizer 时回退 AutoTokenizer。/v1/completions 走 encode(prompt)。chat 且无 chat_template 时用手拼 DSV4 marker 的 _encode_dsv4_messages。 5. **单测** 覆盖单 HTTP 展开多 DP 注册、npu-only medium map、单口折叠命中、DSV4 tokenizer 回退。 ## **3. 资料变更** 不涉及。调度配置项(scheduler_type=kv_cache_affinity、kv_affinity、kv-events-config)沿用现有文档,无新增用户参数。 ## **4. 接口变更** 不涉及。无跨仓协议变更,无客户面 REST/ZMQ 字段变更。kv-conductor /register、/query 语义不变,仅 Coordinator 注册的 DP 数量和打分时读取的 DP 集合按 SGLang 拓扑对齐。 ## **5. 测试结果** **场景:** 800I A3,CANN 9.1.0 B070,MindIE-Motor SGLang nightly,DeepSeek-V4-Flash,2P1D,dp-size=16,DPA,单 HTTP,block_size=128。 **方法:** 同一长前缀(约 4k~6k token)POST /v1/completions 连打;看 Coordinator conductor query response 与 scheduled role=prefill policy=。 **功能验证:** **冷启动第一次 miss,回退负载均衡:**  **第二次起开始命中:**  **性能验证:**  ## **6. CheckList** PR提交人对以下CheckList自检项进行全量自检,自检通过或不涉及,均修改 [ ] 为 [x] [x] 代码注释完备 [x] 正确记录维测日志 [x] 是否有UT用例 [x] 若涉及多线程场景,考虑了并发场景,不存在死锁问题 See merge request: Ascend/MindIE-Motor!1036 | 7 天前 | |
【Feature】新增mindie-motor_standalone_coordinator部署模式 Co-authored-by: zhang980530<zhanghao680@h-partners.com> # message auto-generated for no-merge-commit merge: !910 merge master into master 【Feature】新增mindie-motor_standalone_coordinator部署模式 Created-by: zhang980530 Commit-by: zhang980530 Merged-by: towncharlie Description: ## **1. 合入背景** Fixes [#546](https://gitcode.com/Ascend/MindIE-Motor/issues/546) ## **2. 修改内容** 本 PR 为 MindIE Motor 引入"独立部署(Standalone)Coordinator"能力:通过新增 controller_integration_enabled 配置项,使 Coordinator 可在不依赖 Controller 控制面的场景下独立运行,并新增一套仅含路由信息的"外部部署器(External Deployer)"实例事件协议,将最小拓扑信息转换为内部 InsEventMsg 后复用既有纳管与调度流程。同时补充了 KV conductor 在未配置时的跳过逻辑、裸机部署的配置与启停脚本示例,以及配套的单元测试和端到端冒烟测试。 ## **3. 资料变更** 涉及 MindIE Motor 引入"独立部署(Standalone)Coordinator"能力 ## **4. 接口变更** 不涉及 ## **5. 测试结果**  ------------------------------------------------------------------------------- 用例 C1 文档 SET 1P1D + 模型名自动发现 + readiness 结果 通过 时间 08:47:00 -------------------------------------------------------------------------------- 步骤:POST /instances/refresh body=topology_set_1p1d.json(无 model_name) HTTP 响应 resp_set_1p1d.json {"status":"success","message":"Instance refresh completed", "data":{"event_type":"set","instance_count":2, "timestamp":"2026-09-09T00:47:00.522166+00:00"}} GET /readiness readiness_after_set_1p1d.json {"status":"ok","message":"Coordinator is ok","ready":true} GET /instances instances_after_set_1p1d.json count=2 id=1 prefill model_name=qwen3-8b job_name=external-prefill-1 endpoint id=0 port=8100 id=3 decode model_name=qwen3-8b job_name=external-decode-3 endpoint id=0 port=8200 Coordinator 日志(node/node-37-110.log / coordinator.log) 08:47:00 Refresh instances: event_type=EventType.SET, count=2, instance_ids=[(1, 'prefill'), (3, 'decode')] 08:47:00 Added instance ID 1 (role: prefill, job_name: external-prefill-1) with 1 endpoints to available pool successfully 08:47:00 Added instance ID 3 (role: decode, job_name: external-decode-3) with 1 endpoints to available pool successfully 08:47:00 HBM DP registered (ZMQ+HTTP): instance=vllm-prefill-1 dp=0 replay=tcp://127.0.0.1:6677 08:47:00 Refresh instances done: E=0, P=1, D=1, U=0 判定:省略 model_name 仍按精简协议解析;模型从引擎发现为 qwen3-8b;P/D 齐后 ready=true。 -------------------------------------------------------------------------------- 用例 C2 ADD P2 不启动进程 / DEL P2 不停止进程 结果 通过 时间 08:47:06 -------------------------------------------------------------------------------- ADD 响应 resp_add_p2.json {"status":"success","data":{"event_type":"add","instance_count":1, "timestamp":"2026-09-09T00:47:06.115084+00:00"}} ADD 后实例表 instances_after_add_p2.json count=3 新增 id=2 prefill job_name=external-prefill-2 endpoint id=0 port=8110 DEL 响应 resp_del_p2.json {"status":"success","data":{"event_type":"del","instance_count":1, "timestamp":"2026-09-09T00:47:06.318868+00:00"}} DEL 后实例表 instances_after_del_p2.json count=2(仅 P1+D) 进程核验(当时现场) p2_pid_before=39516 p2_pid_after_add=39516 same=1 p2_pid_after_del=39516 health=200 still_alive=1 日志 08:47:06 Refresh instances: event_type=EventType.ADD, count=1, instance_ids=[(2, 'prefill')] 08:47:06 Added instance ID 2 ... Refresh instances done: E=0, P=2, D=1, U=0 08:47:06 Refresh instances: event_type=EventType.DEL, count=1, instance_ids=[(2, 'prefill')] 08:47:06 Refresh instances done: E=0, P=1, D=1, U=0 判定:add/del 只改 Coordinator 纳管表,不启停 P2 进程。 -------------------------------------------------------------------------------- 用例 C3 Coordinator 重启后必须重放 SET,重放后推理恢复 结果 通过 时间 08:55:58 ~ 08:56:16 -------------------------------------------------------------------------------- 重启后、重放前 readiness_before_replay.json {"status":"ok","message":"Coordinator is ok","ready":false} 重放 SET resp_replay_set.json {"status":"success","data":{"event_type":"set","instance_count":3, "timestamp":"2026-09-09T00:56:06.200305+00:00"}} 重放后 readiness_after_replay.json {"status":"ok","message":"Coordinator is ok","ready":true} 推理 infer_after_replay.json id=cmpl-178891537681151000006c969451 text=" I'm a new user, so I'm not sure what to do. I" usage prompt_tokens=6 total_tokens=22 completion_tokens=16 coordinator_restart.log 08:56:04 Uvicorn running on http://0.0.0.0:19001 08:56:06 Refresh instances: event_type=EventType.SET, count=3, instance_ids=[(1, 'prefill'), (2, 'prefill'), (3, 'decode')] 08:56:06 Added instance ID 1/2/3 ... to available pool successfully 08:56:06 HBM DP registered: vllm-prefill-1 dp=0 ; vllm-prefill-2 dp=0 08:56:06 Refresh instances done: E=0, P=2, D=1, U=0 08:56:06 AUDIT event_type=instance_refresh result=success method=POST path=/instances/refresh 08:56:06 "POST /instances/refresh HTTP/1.1" 200 OK 08:56:06 [Readiness] Coordinator is ready. result=ok_standby instances_status=required_met 08:56:16 Worker 0: metaserver port enabled: 19100 08:56:16 scheduled role=decode instance=3 endpoint=0 08:56:16 scheduled role=prefill instance=1 endpoint=0 req_id=178891537681151000006c969451 policy=kv_cache_affinity matched=0 proposed=1-0 HTTPClientPool address=127.0.0.1:8100 以及 127.0.0.1:8200 判定:重启丢内存拓扑,ready=false;外部重放完整 set 后 ready=true,Layerwise 推理恢复(含 metaserver 19100)。 -------------------------------------------------------------------------------- 用例 C4 Layerwise 1P1D 推理(concurrent_engine_sync) 结果 通过 时间 08:53:22 -------------------------------------------------------------------------------- SET 响应 resp_set_1p1d_concurrent.json {"status":"success","data":{"event_type":"set","instance_count":2}} 推理 POST :19000/v1/completions infer_1p1d_concurrent.json id=cmpl-17889152027183390000b61d2582 model=qwen3-8b text=" I'm a new user, so I'm not sure what to do. I" usage prompt_tokens=6 total_tokens=22 completion_tokens=16 P1 /health 仍为 200 日志 08:53:22 Refresh instances done: E=0, P=1, D=1, U=0 08:53:22 CircuitBreaker cleared all: count=1 08:53:22 scheduled role=prefill req_id=17889152027183390000b61d2582 instance=1 endpoint=0 policy=kv_cache_affinity matched=0 proposed=1-0 判定:精简 SET + concurrent 后,Coordinator 作为推理入口打通 1P1D。 -------------------------------------------------------------------------------- 用例 C5 文档 SET 2P1D,endpoint.id 全部为 0 结果 通过 时间 08:53:42 -------------------------------------------------------------------------------- SET 响应 resp_set_2p1d_concurrent.json instance_count=3 GET /instances instances_2p1d.json id=1 prefill endpoint id=0 port=8100 id=2 prefill endpoint id=0 port=8110 id=3 decode endpoint id=0 port=8200 日志 08:53:42 Refresh instances: event_type=EventType.SET, count=3, instance_ids=[(1, 'prefill'), (2, 'prefill'), (3, 'decode')] 08:53:42 Added instance ID 2 (role: prefill, job_name: external-prefill-2) 08:53:42 HBM DP registered: instance=vllm-prefill-2 dp=0 replay=tcp://127.0.0.1:6677 08:53:42 Refresh instances done: E=0, P=2, D=1, U=0 判定:与昨晚错误 JSON(P2 endpoint.id=1)不同,本次两边都是 id=0、dp=0。 -------------------------------------------------------------------------------- 用例 C6 P1 KV 亲和:热请求 matched=640,查找 key 为 endpoint.id=0 结果 通过 时间 08:53:54 ~ 08:53:56 -------------------------------------------------------------------------------- 请求 affinity_request.json 约 666 prompt tokens,block=128,期望命中 5 块=640。 冷请求 p1_warm.json id=chatcmpl-17889152340817130000091faa74 usage prompt_tokens=666 热请求 p1_hot.json id=chatcmpl-178891523688943700001fe22583 usage prompt_tokens=666 completion_tokens=8 日志 08:53:54 scheduled ... instance=1 endpoint=0 policy=load_balance matched=None req_id=17889152340817130000091faa74 committed=666.0 08:53:56 scheduled ... instance=1 endpoint=0 policy=kv_cache_affinity matched=640 req_id=178891523688943700001fe22583 committed=26.0 proposed=1-0 判定:文档 id=0 与引擎 DP0、调度 DP["0"] 对齐。昨晚 id=1 → DP["1"] → matched=0 的问题本次未复现。 -------------------------------------------------------------------------------- 用例 C7 P2 独有前缀亲和:订 5578 后 matched=384,endpoint=0 结果 通过 时间 08:55:35 ~ 08:55:37 -------------------------------------------------------------------------------- 手工订 conductor(SET 协议带不了 per-instance ZMQ 口) POST :19333/register resp_p2_register.json {"status":"ok"} instance_id=vllm-prefill-2 dp_rank=0 npu=tcp://127.0.0.1:5578 replay=6678 独有前缀 p2_unique_hot.json id=chatcmpl-17889153379236990000289e62a6 usage prompt_tokens=424 total_tokens=432 日志 08:55:35 scheduled instance=2 endpoint=0 policy=load_balance matched=None committed=424.0 req_id=178891533516689100006adc96b5 08:55:37 scheduled instance=2 endpoint=0 policy=kv_cache_affinity matched=384 req_id=17889153379236990000289e62a6 proposed=2-0 判定:查找 key 仍是 "0"(不是 "1"),P2 自己的事件能计分。 附:自动注册时 P1/P2 都订 replay=tcp://127.0.0.1:6677(5577+id,id 都是 0)。 P2 真实口是 5578。这是同机多 P 的 conductor 端口偏移限制,不是 SET id 写错。 -------------------------------------------------------------------------------- 用例 C8 liu意见2:省略 model_name 的精简 SET 结果 通过 时间 09:41:16 -------------------------------------------------------------------------------- liu_omit_model.json {"status":"success","data":{"event_type":"set","instance_count":2}} http=200 liu_instances.json id=1 prefill model_name=qwen3-8b job_name=external-prefill-1 endpoint id=0:8100 id=3 decode model_name=qwen3-8b job_name=external-decode-3 endpoint id=0:8200 -------------------------------------------------------------------------------- 用例 C9 liu意见3:拒绝 decode_colocation,放行 concurrent_engine_sync 结果 通过 时间 09:41:17 -------------------------------------------------------------------------------- liu_reject_colocation.out HTTP 400 External Deployer dispatch_capabilities must be one of: prefill_handoff_decode, concurrent_engine_sync input_value='decode_colocation' liu_allow_concurrent.out HTTP 200 {"status":"success","data":{"event_type":"set","instance_count":2}} -------------------------------------------------------------------------------- 用例 C10 liu意见1:无配置 / 无 Controller 段时不按有 Contro See merge request: Ascend/MindIE-Motor!910 | 22 天前 | |
【Feature】新增mindie-motor_standalone_coordinator部署模式 Co-authored-by: zhang980530<zhanghao680@h-partners.com> # message auto-generated for no-merge-commit merge: !910 merge master into master 【Feature】新增mindie-motor_standalone_coordinator部署模式 Created-by: zhang980530 Commit-by: zhang980530 Merged-by: towncharlie Description: ## **1. 合入背景** Fixes [#546](https://gitcode.com/Ascend/MindIE-Motor/issues/546) ## **2. 修改内容** 本 PR 为 MindIE Motor 引入"独立部署(Standalone)Coordinator"能力:通过新增 controller_integration_enabled 配置项,使 Coordinator 可在不依赖 Controller 控制面的场景下独立运行,并新增一套仅含路由信息的"外部部署器(External Deployer)"实例事件协议,将最小拓扑信息转换为内部 InsEventMsg 后复用既有纳管与调度流程。同时补充了 KV conductor 在未配置时的跳过逻辑、裸机部署的配置与启停脚本示例,以及配套的单元测试和端到端冒烟测试。 ## **3. 资料变更** 涉及 MindIE Motor 引入"独立部署(Standalone)Coordinator"能力 ## **4. 接口变更** 不涉及 ## **5. 测试结果**  ------------------------------------------------------------------------------- 用例 C1 文档 SET 1P1D + 模型名自动发现 + readiness 结果 通过 时间 08:47:00 -------------------------------------------------------------------------------- 步骤:POST /instances/refresh body=topology_set_1p1d.json(无 model_name) HTTP 响应 resp_set_1p1d.json {"status":"success","message":"Instance refresh completed", "data":{"event_type":"set","instance_count":2, "timestamp":"2026-09-09T00:47:00.522166+00:00"}} GET /readiness readiness_after_set_1p1d.json {"status":"ok","message":"Coordinator is ok","ready":true} GET /instances instances_after_set_1p1d.json count=2 id=1 prefill model_name=qwen3-8b job_name=external-prefill-1 endpoint id=0 port=8100 id=3 decode model_name=qwen3-8b job_name=external-decode-3 endpoint id=0 port=8200 Coordinator 日志(node/node-37-110.log / coordinator.log) 08:47:00 Refresh instances: event_type=EventType.SET, count=2, instance_ids=[(1, 'prefill'), (3, 'decode')] 08:47:00 Added instance ID 1 (role: prefill, job_name: external-prefill-1) with 1 endpoints to available pool successfully 08:47:00 Added instance ID 3 (role: decode, job_name: external-decode-3) with 1 endpoints to available pool successfully 08:47:00 HBM DP registered (ZMQ+HTTP): instance=vllm-prefill-1 dp=0 replay=tcp://127.0.0.1:6677 08:47:00 Refresh instances done: E=0, P=1, D=1, U=0 判定:省略 model_name 仍按精简协议解析;模型从引擎发现为 qwen3-8b;P/D 齐后 ready=true。 -------------------------------------------------------------------------------- 用例 C2 ADD P2 不启动进程 / DEL P2 不停止进程 结果 通过 时间 08:47:06 -------------------------------------------------------------------------------- ADD 响应 resp_add_p2.json {"status":"success","data":{"event_type":"add","instance_count":1, "timestamp":"2026-09-09T00:47:06.115084+00:00"}} ADD 后实例表 instances_after_add_p2.json count=3 新增 id=2 prefill job_name=external-prefill-2 endpoint id=0 port=8110 DEL 响应 resp_del_p2.json {"status":"success","data":{"event_type":"del","instance_count":1, "timestamp":"2026-09-09T00:47:06.318868+00:00"}} DEL 后实例表 instances_after_del_p2.json count=2(仅 P1+D) 进程核验(当时现场) p2_pid_before=39516 p2_pid_after_add=39516 same=1 p2_pid_after_del=39516 health=200 still_alive=1 日志 08:47:06 Refresh instances: event_type=EventType.ADD, count=1, instance_ids=[(2, 'prefill')] 08:47:06 Added instance ID 2 ... Refresh instances done: E=0, P=2, D=1, U=0 08:47:06 Refresh instances: event_type=EventType.DEL, count=1, instance_ids=[(2, 'prefill')] 08:47:06 Refresh instances done: E=0, P=1, D=1, U=0 判定:add/del 只改 Coordinator 纳管表,不启停 P2 进程。 -------------------------------------------------------------------------------- 用例 C3 Coordinator 重启后必须重放 SET,重放后推理恢复 结果 通过 时间 08:55:58 ~ 08:56:16 -------------------------------------------------------------------------------- 重启后、重放前 readiness_before_replay.json {"status":"ok","message":"Coordinator is ok","ready":false} 重放 SET resp_replay_set.json {"status":"success","data":{"event_type":"set","instance_count":3, "timestamp":"2026-09-09T00:56:06.200305+00:00"}} 重放后 readiness_after_replay.json {"status":"ok","message":"Coordinator is ok","ready":true} 推理 infer_after_replay.json id=cmpl-178891537681151000006c969451 text=" I'm a new user, so I'm not sure what to do. I" usage prompt_tokens=6 total_tokens=22 completion_tokens=16 coordinator_restart.log 08:56:04 Uvicorn running on http://0.0.0.0:19001 08:56:06 Refresh instances: event_type=EventType.SET, count=3, instance_ids=[(1, 'prefill'), (2, 'prefill'), (3, 'decode')] 08:56:06 Added instance ID 1/2/3 ... to available pool successfully 08:56:06 HBM DP registered: vllm-prefill-1 dp=0 ; vllm-prefill-2 dp=0 08:56:06 Refresh instances done: E=0, P=2, D=1, U=0 08:56:06 AUDIT event_type=instance_refresh result=success method=POST path=/instances/refresh 08:56:06 "POST /instances/refresh HTTP/1.1" 200 OK 08:56:06 [Readiness] Coordinator is ready. result=ok_standby instances_status=required_met 08:56:16 Worker 0: metaserver port enabled: 19100 08:56:16 scheduled role=decode instance=3 endpoint=0 08:56:16 scheduled role=prefill instance=1 endpoint=0 req_id=178891537681151000006c969451 policy=kv_cache_affinity matched=0 proposed=1-0 HTTPClientPool address=127.0.0.1:8100 以及 127.0.0.1:8200 判定:重启丢内存拓扑,ready=false;外部重放完整 set 后 ready=true,Layerwise 推理恢复(含 metaserver 19100)。 -------------------------------------------------------------------------------- 用例 C4 Layerwise 1P1D 推理(concurrent_engine_sync) 结果 通过 时间 08:53:22 -------------------------------------------------------------------------------- SET 响应 resp_set_1p1d_concurrent.json {"status":"success","data":{"event_type":"set","instance_count":2}} 推理 POST :19000/v1/completions infer_1p1d_concurrent.json id=cmpl-17889152027183390000b61d2582 model=qwen3-8b text=" I'm a new user, so I'm not sure what to do. I" usage prompt_tokens=6 total_tokens=22 completion_tokens=16 P1 /health 仍为 200 日志 08:53:22 Refresh instances done: E=0, P=1, D=1, U=0 08:53:22 CircuitBreaker cleared all: count=1 08:53:22 scheduled role=prefill req_id=17889152027183390000b61d2582 instance=1 endpoint=0 policy=kv_cache_affinity matched=0 proposed=1-0 判定:精简 SET + concurrent 后,Coordinator 作为推理入口打通 1P1D。 -------------------------------------------------------------------------------- 用例 C5 文档 SET 2P1D,endpoint.id 全部为 0 结果 通过 时间 08:53:42 -------------------------------------------------------------------------------- SET 响应 resp_set_2p1d_concurrent.json instance_count=3 GET /instances instances_2p1d.json id=1 prefill endpoint id=0 port=8100 id=2 prefill endpoint id=0 port=8110 id=3 decode endpoint id=0 port=8200 日志 08:53:42 Refresh instances: event_type=EventType.SET, count=3, instance_ids=[(1, 'prefill'), (2, 'prefill'), (3, 'decode')] 08:53:42 Added instance ID 2 (role: prefill, job_name: external-prefill-2) 08:53:42 HBM DP registered: instance=vllm-prefill-2 dp=0 replay=tcp://127.0.0.1:6677 08:53:42 Refresh instances done: E=0, P=2, D=1, U=0 判定:与昨晚错误 JSON(P2 endpoint.id=1)不同,本次两边都是 id=0、dp=0。 -------------------------------------------------------------------------------- 用例 C6 P1 KV 亲和:热请求 matched=640,查找 key 为 endpoint.id=0 结果 通过 时间 08:53:54 ~ 08:53:56 -------------------------------------------------------------------------------- 请求 affinity_request.json 约 666 prompt tokens,block=128,期望命中 5 块=640。 冷请求 p1_warm.json id=chatcmpl-17889152340817130000091faa74 usage prompt_tokens=666 热请求 p1_hot.json id=chatcmpl-178891523688943700001fe22583 usage prompt_tokens=666 completion_tokens=8 日志 08:53:54 scheduled ... instance=1 endpoint=0 policy=load_balance matched=None req_id=17889152340817130000091faa74 committed=666.0 08:53:56 scheduled ... instance=1 endpoint=0 policy=kv_cache_affinity matched=640 req_id=178891523688943700001fe22583 committed=26.0 proposed=1-0 判定:文档 id=0 与引擎 DP0、调度 DP["0"] 对齐。昨晚 id=1 → DP["1"] → matched=0 的问题本次未复现。 -------------------------------------------------------------------------------- 用例 C7 P2 独有前缀亲和:订 5578 后 matched=384,endpoint=0 结果 通过 时间 08:55:35 ~ 08:55:37 -------------------------------------------------------------------------------- 手工订 conductor(SET 协议带不了 per-instance ZMQ 口) POST :19333/register resp_p2_register.json {"status":"ok"} instance_id=vllm-prefill-2 dp_rank=0 npu=tcp://127.0.0.1:5578 replay=6678 独有前缀 p2_unique_hot.json id=chatcmpl-17889153379236990000289e62a6 usage prompt_tokens=424 total_tokens=432 日志 08:55:35 scheduled instance=2 endpoint=0 policy=load_balance matched=None committed=424.0 req_id=178891533516689100006adc96b5 08:55:37 scheduled instance=2 endpoint=0 policy=kv_cache_affinity matched=384 req_id=17889153379236990000289e62a6 proposed=2-0 判定:查找 key 仍是 "0"(不是 "1"),P2 自己的事件能计分。 附:自动注册时 P1/P2 都订 replay=tcp://127.0.0.1:6677(5577+id,id 都是 0)。 P2 真实口是 5578。这是同机多 P 的 conductor 端口偏移限制,不是 SET id 写错。 -------------------------------------------------------------------------------- 用例 C8 liu意见2:省略 model_name 的精简 SET 结果 通过 时间 09:41:16 -------------------------------------------------------------------------------- liu_omit_model.json {"status":"success","data":{"event_type":"set","instance_count":2}} http=200 liu_instances.json id=1 prefill model_name=qwen3-8b job_name=external-prefill-1 endpoint id=0:8100 id=3 decode model_name=qwen3-8b job_name=external-decode-3 endpoint id=0:8200 -------------------------------------------------------------------------------- 用例 C9 liu意见3:拒绝 decode_colocation,放行 concurrent_engine_sync 结果 通过 时间 09:41:17 -------------------------------------------------------------------------------- liu_reject_colocation.out HTTP 400 External Deployer dispatch_capabilities must be one of: prefill_handoff_decode, concurrent_engine_sync input_value='decode_colocation' liu_allow_concurrent.out HTTP 200 {"status":"success","data":{"event_type":"set","instance_count":2}} -------------------------------------------------------------------------------- 用例 C10 liu意见1:无配置 / 无 Controller 段时不按有 Contro See merge request: Ascend/MindIE-Motor!910 | 22 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 7 个月前 | ||
| 7 天前 | ||
| 22 天前 | ||
| 22 天前 |