已关闭
[Bug-Report|缺陷反馈]: DeepSeek-V3.1-Terminus 量化报错 #307
贾维斯Echo创建于  2025年10月2日关闭于  2025年10月20日
贾维斯Echo
贾维斯Echo
2025年10月2日 创建

Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.

Describe the current behavior / 问题描述 (Mandatory / 必填)

量化过程如下:

source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:False
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
conda create -y -n msit_env python=3.10
conda  activate msit_env
pip3 install attrs cython 'numpy>=1.19.2,<=1.24.0' decorator sympy cffi pyyaml pathlib2 psutil protobuf==3.20.0 scipy requests absl-py --user
git clone https://gitcode.com/Ascend/msit.git
cd /mnt/data/msit/msmodelslim
bash install.sh 
pip install torch==2.1.0 torch_npu==7.1.0 --extra-index-url "https://download.pytorch.org/whl/cpu/" --extra-index-url "https://mirrors.huaweicloud.com/ascend/repos/pypi/"
pip install transformers==4.48.2

Environment / 环境信息 (Mandatory / 必填)

python 环境:3.10
昇腾硬件型:910B2

机器配置如下:


root@pm-a206 ~]# npu-smi info
+------------------------------------------------------------------------------------------------+
| npu-smi 24.1.0.3                 Version: 24.1.0.3                                             |
+---------------------------+---------------+----------------------------------------------------+
| NPU   Name                | Health        | Power(W)    Temp(C)           Hugepages-Usage(page)|
| Chip                      | Bus-Id        | AICore(%)   Memory-Usage(MB)  HBM-Usage(MB)        |
+===========================+===============+====================================================+
| 0     910B2               | OK            | 94.5        40                0    / 0             |
| 0                         | 0000:C1:00.0  | 0           0    / 0          3390 / 65536         |
+===========================+===============+====================================================+
| 1     910B2               | OK            | 93.2        40                0    / 0             |
| 0                         | 0000:01:00.0  | 0           0    / 0          3389 / 65536         |
+===========================+===============+====================================================+
| 2     910B2               | OK            | 95.6        40                0    / 0             |
| 0                         | 0000:C2:00.0  | 0           0    / 0          3393 / 65536         |
+===========================+===============+====================================================+
| 3     910B2               | OK            | 93.3        41                0    / 0             |
| 0                         | 0000:02:00.0  | 0           0    / 0          3389 / 65536         |
+===========================+===============+====================================================+
| 4     910B2               | OK            | 95.8        40                0    / 0             |
| 0                         | 0000:81:00.0  | 0           0    / 0          3389 / 65536         |
+===========================+===============+====================================================+
| 5     910B2               | OK            | 97.0        41                0    / 0             |
| 0                         | 0000:41:00.0  | 0           0    / 0          3392 / 65536         |
+===========================+===============+====================================================+
| 6     910B2               | OK            | 94.4        40                0    / 0             |
| 0                         | 0000:82:00.0  | 0           0    / 0          3390 / 65536         |
+===========================+===============+====================================================+
| 7     910B2               | OK            | 99.2        42                0    / 0             |
| 0                         | 0000:42:00.0  | 0           0    / 0          3390 / 65536         |
+===========================+===============+====================================================+
+---------------------------+---------------+----------------------------------------------------+
| NPU     Chip              | Process id    | Process name             | Process memory(MB)      |
+===========================+===============+====================================================+
| No running processes found in NPU 0                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 1                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 2                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 3                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 4                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 5                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 6                                                            |
+===========================+===============+====================================================+
| No running processes found in NPU 7                                                            |
+===========================+===============+====================================================+

Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)

模型我是下载的这里的:https://modelscope.cn/models/unsloth/DeepSeek-V3.1-Terminus-BF16

启动命令如下:

cd /mnt/data/msit/msmodelslim/example/DeepSeek
nohup python3 quant_deepseek_w8a8.py \
--model_path /mnt/nvme0n1/model/DeepSeek-V3.1-Terminus-BF16/DeepSeek-V3.1-Terminus-BF16 \
--save_path /mnt/data/model/DeepSeek-V3.1-Terminus-w8a8 \
--batch_size 4 \
--anti_dataset ./anti_prompt_50_v3_1.json \
--calib_dataset ./calib_prompt_50_v3_1.json \
--anti_method m4 \
--quant_mtp mix \
--rot > /mnt/data/log/quant_w8a8.log 2>&1 &

Describe the expected behavior / 预期结果 (Mandatory / 必填)

应该正常量化成功啊,但是始终报错

错误日志如下:

[root@pm-a206 ~]# cat /mnt/data/log/quant_w8a8.log
nohup: ignoring input
/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/site-packages/torch_npu/contrib/transfer_to_npu.py:292: ImportWarning: 
    *************************************************************************************************************
    The torch.Tensor.cuda and torch.nn.Module.cuda are replaced with torch.Tensor.npu and torch.nn.Module.npu now..
    The torch.cuda.DoubleTensor is replaced with torch.npu.FloatTensor cause the double type is not supported now..
    The backend in torch.distributed.init_process_group set to hccl now..
    The torch.cuda.* and torch.cuda.amp.* are replaced with torch.npu.* and torch.npu.amp.* now..
    The device parameters have been replaced with npu in the function below:
    torch.logspace, torch.randint, torch.hann_window, torch.rand, torch.full_like, torch.ones_like, torch.rand_like, torch.randperm, torch.arange, torch.frombuffer, torch.normal, torch._empty_per_channel_affine_quantized, torch.empty_strided, torch.empty_like, torch.scalar_tensor, torch.tril_indices, torch.bartlett_window, torch.ones, torch.sparse_coo_tensor, torch.randn, torch.kaiser_window, torch.tensor, torch.triu_indices, torch.as_tensor, torch.zeros, torch.randint_like, torch.full, torch.eye, torch._sparse_csr_tensor_unsafe, torch.empty, torch._sparse_coo_tensor_unsafe, torch.blackman_window, torch.zeros_like, torch.range, torch.sparse_csr_tensor, torch.randn_like, torch.from_file, torch._cudnn_init_dropout_state, torch._empty_affine_quantized, torch.linspace, torch.hamming_window, torch.empty_quantized, torch._pin_memory, torch.autocast, torch.load, torch.Generator, torch.set_default_device, torch.Tensor.new_empty, torch.Tensor.new_empty_strided, torch.Tensor.new_full, torch.Tensor.new_ones, torch.Tensor.new_tensor, torch.Tensor.new_zeros, torch.Tensor.to, torch.Tensor.pin_memory, torch.nn.Module.to, torch.nn.Module.to_empty
    *************************************************************************************************************
    
  warnings.warn(msg, ImportWarning)
/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/site-packages/torch_npu/contrib/transfer_to_npu.py:247: RuntimeWarning: torch.jit.script and torch.jit.script_method will be disabled by transfer_to_npu, which currently does not support them, if you need to enable them, please do not use transfer_to_npu.
  warnings.warn(msg, RuntimeWarning)
Total Process:   0%|          | 0/5 [00:00<?, ?it/s]2025-10-01 23:19:31,037 - msmodelslim - INFO - write directory exists, write file to directory '/mnt/data/model/DeepSeek-V3.1-Terminus-w8a8'
Loading checkpoint shards: 100%|██████████| 163/163 [02:21<00:00,  1.15it/s]
Some weights of the model checkpoint at /mnt/nvme0n1/model/DeepSeek-V3.1-Terminus-BF16/DeepSeek-V3.1-Terminus-BF16 were not used when initializing DeepseekV3ForCausalLM: {'model.layers.61.embed_tokens.weight', 'model.layers.61.eh_proj.weight', 'model.layers.61.shared_head.head.weight', 'model.layers.61.hnorm.weight', 'model.layers.61.shared_head.norm.weight', 'model.layers.61.enorm.weight'}
- This IS expected if you are initializing DeepseekV3ForCausalLM from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model).
- This IS NOT expected if you are initializing DeepseekV3ForCausalLM from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model).
Some parameters are on the meta device because they were offloaded to the cpu.
54it [1:19:04, 87.87s/it]
Traceback (most recent call last):
  File "/mnt/data/msit/msmodelslim/example/DeepSeek/quant_deepseek_w8a8.py", line 317, in <module>
    main()
  File "/mnt/data/msit/msmodelslim/example/DeepSeek/quant_deepseek_w8a8.py", line 211, in main
    rot_model(model)
  File "/mnt/data/msit/msmodelslim/example/common/rot_utils/rot_ds.py", line 633, in rot_model
    fuse_layer_norms(model, fuse_kv_ln=rotate_kv)
  File "/mnt/data/msit/msmodelslim/example/common/rot_utils/rot_ds.py", line 198, in fuse_layer_norms
    fuse_layer_norms_module(layer, model.config.hidden_size, fuse_kv_ln, idx=i, model=model)
  File "/mnt/data/msit/msmodelslim/example/common/rot_utils/rot_ds.py", line 92, in fuse_layer_norms_module
    with ResListToRelease(*prepare_list):
  File "/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/site-packages/ascend_utils/common/utils.py", line 125, in __exit__
    res.__exit__(exc_type, exc_val, exc_tb)
  File "/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/site-packages/msmodelslim/pytorch/llm_ptq/accelerate_adapter/hook_adapter.py", line 60, in __exit__
    hook.post_forward(self.module, *[torch.zeros([1])])
  File "/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/site-packages/msmodelslim/pytorch/llm_ptq/accelerate_adapter/hook_adapter.py", line 246, in post_forward
    self.update_weights_map(module, force_update=self.post_force, recurse=self.post_recurse)
  File "/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/site-packages/msmodelslim/pytorch/llm_ptq/accelerate_adapter/hook_adapter.py", line 208, in update_weights_map
    dataset[weights_map.prefix + key] = item.clone().detach().cpu()
RuntimeError: [enforce fail at alloc_cpu.cpp:119] err == 0. DefaultCPUAllocator: can't allocate memory: you tried to allocate 75497472 bytes. Error code 12 (Cannot allocate memory)
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
[ERROR] TBE Subprocess[task_distribute] raise error[], main process disappeared!
/mnt/nvme0n1/app_data/conda/envs/msit_env/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning: resource_tracker: There appear to be 30 leaked semaphore objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '

Special notes for this issue/备注 (Optional / 选填)

likedislike
贾维斯Echo贾维斯Echo
2025年10月6日 修改标题为 “[Bug-Report|缺陷反馈]: DeepSeek-V3.1-Terminus 量化报错,”,原标题为“[Bug-Report|缺陷反馈]: DeepSeek-V3.1-Terminus 量化失败,试了切换各种torch 版本始终报错”
贾维斯Echo贾维斯Echo
2025年10月6日 修改标题为 “[Bug-Report|缺陷反馈]: DeepSeek-V3.1-Terminus 量化报错”,原标题为“[Bug-Report|缺陷反馈]: DeepSeek-V3.1-Terminus 量化报错,”
贾维斯Echo
贾维斯Echo
2025年10月6日 评论:

使用原版模型DeepSeek-V3.1-Terminus量化成功了,但是依旧最后一步报错如下:

Traceback (most recent call last):
  File "/mnt/data/msit/msmodelslim/example/DeepSeek/quant_deepseek_w8a8.py", line 317, in <module>
    main()
  File "/mnt/data/msit/msmodelslim/example/DeepSeek/quant_deepseek_w8a8.py", line 306, in main
    post_process_mtp_quant(save_path)
  File "/mnt/data/msit/msmodelslim/example/DeepSeek/mtp_quant_module.py", line 409, in post_process_mtp_quant
    os.chmod(tensor_path, READ_WRITE_PERMISSION, follow_symlinks=False)
NotImplementedError: chmod: follow_symlinks unavailable on this platform
[ERROR] 2025-10-06-04:50:16 (PID:3239210, Device:0, RankID:-1) ERR99999 UNKNOWN applicaiton exception
Total Process:  80%|████████  | 4/5 [6:11:37<1:32:54, 5574.28s/it]
likedislike
ninth99
ninth99
2025年10月9日 评论:

这个错误是操作系统的权限问题吧。

likedislike
贾维斯Echo
贾维斯Echo
2025年10月9日 评论:

这个错误是操作系统的权限问题吧。

不是,麒麟v10 系统好像不支持 follow_symlinks=False 这个参数,去掉就可以了

likedislike
anreywmh成员
2025年10月14日 评论:

尊敬的用户,您好!由于工单长时间未收到您的响应,现暂为关闭。感谢您对 msModelSlim 开源生态的贡献!后续如有疑问,欢迎随时重新提问。

likedislike
Aanreywmh成员
2025年10月20日 issue状态由 TODO 改变为 DONE
Aanreywmh成员
2025年10月20日 关闭了 issue