https://gitcode.com/Ascend/msmodeling/blob/master/docs/zh/user_guide/msmodeling_throughput_optimizer_user_guide.md#3-结果说明
3 结果说明 命令执行成功后,终端会先打印输入配置和最优配置摘要,随后展示候选并行配置表。输出指标包括吞吐量、TTFT、TPOT、并发度,以及模式相关字段(如 QPS 或 PD 配比)。例如:
Top 4 Aggregation Configurations: +-----+----------------------+-----------+-----------+-------------+-------------+--------------------+------------+ | Top | Throughput (token/s) | TTFT (ms) | TPOT (ms) | concurrency | num_devices | parallel | batch_size | +-----+----------------------+-----------+-----------+-------------+-------------+--------------------+------------+ | 1 | 2888.45 | 16032.05 | 49.90 | 175 | 8 | TP=8 | PP=1 | DP=1 | 175 | | 2 | 2013.49 | 22512.86 | 49.56 | 130 | 8 | TP=4 | PP=1 | DP=2 | 65 | | 3 | 1140.23 | 25817.73 | 49.44 | 76 | 8 | TP=2 | PP=1 | DP=4 | 19 | | 4 | 549.89 | 14214.54 | 48.72 | 32 | 8 | TP=1 | PP=1 | DP=8 | 4 | +-----+----------------------+-----------+-----------+-------------+-------------+--------------------+------------+
文档中描述了多个使用场景,但只给了这一个结果示例,并且这一段前就是2.3PD配比场景,顺序阅读下来很可能以为是PD配比场景的输出(容易误导读者),但这个输出有没有PD配比信息,读者会懵逼,会思考是没写全吗?这个到底是哪个示例的输出? 仔细阅读文档后我发现这个输出可以和2.1 PD 混部场景的示例输入对上,所以希望能加上对应说明,方便阅读,增加易用性。
可以改成: 例如:2.1 PD 混部场景示例执行成功后,可以得到以下配置表。 或者参考 模型推理性能仿真 使用指南,每种场景都给了输入和输出。
Welcome to join the community and thank you for your contribution 🎉!
欢迎加入社区,感谢您对社区的贡献 🎉!
/label add triaged
Documentation Location (Multiple document links can be specified) | 文档位置(可指定多个文档链接)
https://gitcode.com/Ascend/msmodeling/blob/master/docs/zh/user_guide/msmodeling_throughput_optimizer_user_guide.md#3-结果说明
Current Content Description | 文档问题描述
3 结果说明
命令执行成功后,终端会先打印输入配置和最优配置摘要,随后展示候选并行配置表。输出指标包括吞吐量、TTFT、TPOT、并发度,以及模式相关字段(如 QPS 或 PD 配比)。例如:
Input Configuration:
Model: Qwen/Qwen3-32B
Quantize Linear action: W8A8_DYNAMIC
Quantize Attention action: disabled
Devices: 8 TEST_DEVICE
TTFT Limits: None ms
TPOT Limits: 50.0 ms
Overall Best Configuration:
Best Throughput: 2888.45 tokens/s
TTFT: 16032.05 ms
TPOT: 49.90 ms
Top 4 Aggregation Configurations:
+-----+----------------------+-----------+-----------+-------------+-------------+--------------------+------------+
| Top | Throughput (token/s) | TTFT (ms) | TPOT (ms) | concurrency | num_devices | parallel | batch_size |
+-----+----------------------+-----------+-----------+-------------+-------------+--------------------+------------+
| 1 | 2888.45 | 16032.05 | 49.90 | 175 | 8 | TP=8 | PP=1 | DP=1 | 175 |
| 2 | 2013.49 | 22512.86 | 49.56 | 130 | 8 | TP=4 | PP=1 | DP=2 | 65 |
| 3 | 1140.23 | 25817.73 | 49.44 | 76 | 8 | TP=2 | PP=1 | DP=4 | 19 |
| 4 | 549.89 | 14214.54 | 48.72 | 32 | 8 | TP=1 | PP=1 | DP=8 | 4 |
+-----+----------------------+-----------+-----------+-------------+-------------+--------------------+------------+
文档中描述了多个使用场景,但只给了这一个结果示例,并且这一段前就是2.3PD配比场景,顺序阅读下来很可能以为是PD配比场景的输出(容易误导读者),但这个输出有没有PD配比信息,读者会懵逼,会思考是没写全吗?这个到底是哪个示例的输出?
仔细阅读文档后我发现这个输出可以和2.1 PD 混部场景的示例输入对上,所以希望能加上对应说明,方便阅读,增加易用性。
Modification Suggestion | 修改建议
可以改成:
例如:2.1 PD 混部场景示例执行成功后,可以得到以下配置表。
或者参考 模型推理性能仿真 使用指南,每种场景都给了输入和输出。
Welcome to join the community and thank you for your contribution 🎉!
欢迎加入社区,感谢您对社区的贡献 🎉!