Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.
使用DeekSeek-V3.2 fp8权重转int8会爆显存,报错信息如下
910B3 64GB
msmodelslim quant \ --model_path ${model_path} \ --save_path ${save_path} \ --model_type DeepSeek-V3.2-Exp \ --quant_type w8a8 \ --trust_remote_code True
设置了环境变量export PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:128
export PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:128
正常跑通,不报错
2025-12-02 17:49:40,584 - msmodelslim.core.runner.layer_wise_runner - ERROR - Error processing subgraph model.layers.5.post_attention_layernorm: NPU out of memory. Tried to allocate 33.03 GiB (NPU 0; 60.96 GiB total capacity; 35.89 GiB already allocated; 35.89 GiB current active; 22.77 GiB free; 37.41 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.
layer 3 4 5 6报错
您好 请使用br_release_MindStudio_8.3.0_20261231这个分支版本
我看有个pr 4756 ,是否可以使用这个
也是可以的,master已修复这个问题
ok, 验证没有问题了,多谢
Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.
Describe the current behavior / 问题描述 (Mandatory / 必填)
使用DeekSeek-V3.2 fp8权重转int8会爆显存,报错信息如下
Environment / 环境信息 (Mandatory / 必填)
910B3 64GB
Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)
msmodelslim quant \ --model_path ${model_path} \ --save_path ${save_path} \ --model_type DeepSeek-V3.2-Exp \ --quant_type w8a8 \ --trust_remote_code True设置了环境变量
export PYTORCH_NPU_ALLOC_CONF=max_split_size_mb:128Describe the expected behavior / 预期结果 (Mandatory / 必填)
正常跑通,不报错
Related log / screenshot / 日志 / 截图 (Mandatory / 必填)
layer 3 4 5 6报错
Special notes for this issue/备注 (Optional / 选填)