已关闭
[Bug-Report|缺陷反馈]: DeepSeek-V3.2 910B3 一键量化爆显存 #353
flyrae000创建于  2025年12月2日关闭于  2025年12月5日
flyrae000
2025年12月2日 创建

Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.

Describe the current behavior / 问题描述 (Mandatory / 必填)

使用DeekSeek-V3.2 fp8权重转int8会爆显存,报错信息如下

Environment / 环境信息 (Mandatory / 必填)

910B3 64GB

Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)

msmodelslim quant \
 --model_path ${model_path} \
 --save_path ${save_path} \
 --model_type DeepSeek-V3.2-Exp \
 --quant_type w8a8 \
 --trust_remote_code True

设置了环境变量export PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:128

Describe the expected behavior / 预期结果 (Mandatory / 必填)

正常跑通,不报错

2025-12-02 17:49:40,584 - msmodelslim.core.runner.layer_wise_runner - ERROR - Error processing subgraph model.layers.5.post_attention_layernorm: NPU out of memory. Tried to allocate 33.03 GiB (NPU 0; 60.96 GiB total capacity; 35.89 GiB already allocated; 35.89 GiB current active; 22.77 GiB free; 37.41 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.

layer 3 4 5 6报错

Special notes for this issue/备注 (Optional / 选填)

likedislike
anreywmh成员
2025年12月3日 评论:

您好 请使用br_release_MindStudio_8.3.0_20261231这个分支版本

likedislike
flyrae000
2025年12月3日 评论:

我看有个pr 4756 ,是否可以使用这个

likedislike
anreywmh成员
2025年12月4日 评论:

也是可以的,master已修复这个问题

likedislike
flyrae000
2025年12月5日 评论:

ok, 验证没有问题了,多谢

likedislike
Aanreywmh成员
2025年12月5日 issue状态由 TODO 改变为 DONE
Aanreywmh成员
2025年12月5日 关闭了 issue
Aanreywmh成员
2025年12月24日 关联了pull request:【msmodelslim】【bugfix】修复dpskv32 w8a8最佳实践yaml以及iter smooth bugfix
Aanreywmh成员
2025年12月24日 删除了关联的pull request:【msmodelslim】【bugfix】修复dpskv32 w8a8最佳实践yaml以及iter smooth bugfix
Aanreywmh成员
2025年12月24日 关联了pull request:【msmodelslim】【bugfix】修复dpskv32 w8a8最佳实践yaml以及iter smooth bugfix
AAiBoGe成员
1月6日 关联了看板:msit
ascend-robotascend-robot成员
1月22日 添加了label:resolved