您好,请使用800I A2服务器进行量化,mindie 310P不支持部署量化权重
https://www.hiascend.com/software/mindie/modellist?type=多模态理解模型列表


我是310P3的24GB显存的卡,用800I-A2量化好的在310P3上使用报如下错误:
环境:300I-Duo-Mindie2.2.RC1镜像
运行命令:
export ASCEND_RT_VISIBLE_DEVICES=0
export MINDIE_LOG_TO_STDOUT=1
export MINDIE_LOG_TO_FILE=1
export MINDIE_LOG_LEVEL=info
export LD_LIBRARY_PATH=/usr/local/Ascend/atb-models/lib:$LD_LIBRARY_PATH
bash /usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.sh --model_path /workspace/llm-models/Qwen2.5-VL-7B-W8A8 --input_image /workspace/data/test.png --input_text "please explain this image"
错误日志:
见附件4c8ba50f59694e05b3fedbeb879f444a.txt


您好,请使用800I A2服务器进行量化,mindie 310P不支持部署量化权重
https://www.hiascend.com/software/mindie/modellist?type=多模态理解模型列表
310P3支持多模态量化模型的部署吗?试过大语言模型的量化是可以跑的


您可以查看MindIE的官方模型支持列表,没写就是不支持的
https://www.hiascend.com/software/mindie/modellist?type=多模态理解模型列表


Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.
Describe the current behavior / 问题描述 (Mandatory / 必填)
直接git clone https://gitcode.com/Ascend/msit.git
按安装步骤安装 qwen_vl_utils和transformers以及msmodelslim
基于310P的NPU量化Qwen2.5-VL-7B报错,换成cpu出现
Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_outlier.py", line 409, in process
self._process()
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_outlier.py", line 530, in _process
self.model(*data)
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py", line 1795, in forward
image_embeds = self.visual(pixel_values, grid_thw=image_grid_thw)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py", line 518, in forward
hidden_states = self.patch_embed(hidden_states)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py", line 111, in forward
hidden_states = self.proj(hidden_states.to(dtype=target_dtype)).view(-1, self.embed_dim)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/conv.py", line 610, in forward
return self._conv_forward(input, self.weight, self.bias)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/conv.py", line 605, in _conv_forward
return F.conv3d(
^^^^^^^^^
RuntimeError: "compute_columns3d" not implemented for 'Half'
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/workspace/msit-master/msmodelslim/example/multimodal_vlm/Qwen2.5-VL/quant_qwen2_5vl.py", line 120, in
anti_outlier.process()
File "/usr/local/lib64/python3.11/site-packages/torch/utils/_contextlib.py", line 115, in decorate_context
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_outlier.py", line 411, in process
raise Exception("Please check your config, model and input!", e) from e
Exception: ('Please check your config, model and input!', RuntimeError('"compute_columns3d" not implemented for 'Half''))
[ERROR] 2025-12-16-19:56:57 (PID:57535, Device:-1, RankID:-1) ERR99999 UNKNOWN application exception
Environment / 环境信息 (Mandatory / 必填)
driver:
24.1.rc2
cann:
8.2.RC2
**msit: **
commit 602b237d5db01df5447d78485d0cf7eebf17ffcf (HEAD -> master, origin/master, origin/HEAD)
Merge: 953922c28 44f78a735
Author: Secluded_Ocean tangchuxiao0709@qq.com
Date: Tue Dec 16 19:54:47 2025 +0800
python:2.1.0
transformers:4.49.0
Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)
python quant_qwen2_5vl.py --model_path /workspace/llm-models/Qwen2.5-VL-7B-Instruct --calib_images ../calibImages --save_directory /workspace/Qwen2.5-VL-7B-W8A8 --w_bit 8 --a_bit 8 --device_type npu --trust_remote_code True --anti_method m2 --mindie_format
Describe the expected behavior / 预期结果 (Mandatory / 必填)
预期正常量化并在mindie2.2.RC1中正常使用
Related log / screenshot / 日志 / 截图 (Mandatory / 必填)
/usr/local/lib64/python3.11/site-packages/torch_npu/contrib/transfer_to_npu.py:302: ImportWarning:
*************************************************************************************************************
The torch.Tensor.cuda and torch.nn.Module.cuda are replaced with torch.Tensor.npu and torch.nn.Module.npu now..
The torch.cuda.DoubleTensor is replaced with torch.npu.FloatTensor cause the double type is not supported now..
The backend in torch.distributed.init_process_group set to hccl now..
The torch.cuda.* and torch.cuda.amp.* are replaced with torch.npu.* and torch.npu.amp.* now..
The device parameters have been replaced with npu in the function below:
torch.logspace, torch.randint, torch.hann_window, torch.rand, torch.full_like, torch.ones_like, torch.rand_like, torch.randperm, torch.arange, torch.frombuffer, torch.normal, torch._empty_per_channel_affine_quantized, torch.empty_strided, torch.empty_like, torch.scalar_tensor, torch.tril_indices, torch.bartlett_window, torch.ones, torch.sparse_coo_tensor, torch.randn, torch.kaiser_window, torch.tensor, torch.triu_indices, torch.as_tensor, torch.zeros, torch.randint_like, torch.full, torch.eye, torch._sparse_csr_tensor_unsafe, torch.empty, torch._sparse_coo_tensor_unsafe, torch.blackman_window, torch.zeros_like, torch.range, torch.sparse_csr_tensor, torch.randn_like, torch.from_file, torch._cudnn_init_dropout_state, torch._empty_affine_quantized, torch.linspace, torch.hamming_window, torch.empty_quantized, torch._pin_memory, torch.autocast, torch.load, torch.set_default_device, torch.Tensor.new_empty, torch.Tensor.new_empty_strided, torch.Tensor.new_full, torch.Tensor.new_ones, torch.Tensor.new_tensor, torch.Tensor.new_zeros, torch.Tensor.to, torch.Tensor.pin_memory, torch.nn.Module.to, torch.nn.Module.to_empty
*************************************************************************************************************
warnings.warn(msg, ImportWarning)
/usr/local/lib64/python3.11/site-packages/torch_npu/contrib/transfer_to_npu.py:257: RuntimeWarning: torch.jit.script and torch.jit.script_method will be disabled by transfer_to_npu, which currently does not support them, if you need to enable them, please do not use transfer_to_npu.
warnings.warn(msg, RuntimeWarning)
2025-12-16 20:46:19,139 - msmodelslim - WARNING - write directory not exists, creating directory '/workspace/Qwen2.5-VL-7B-W8A8'
Loading checkpoint shards: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████| 5/5 [00:36<00:00, 7.24s/it]
.Using a slow image processor as
use_fastis unset and a slow processor was saved with this model.use_fast=Truewill be the default behavior in v4.52, even if the model was saved with a slow processor. This will result in minor differences in outputs. You'll still be able to use a slow processor withuse_fast=False.2025-12-16 20:47:45,458 - msmodelslim - WARNING - Not all elements in calib_data are torch.Tensor, please make sure that the model can run with model((calib_data[0]))
2025-12-16 20:47:45,459 - msmodelslim - WARNING - Not all elements in calib_data are torch.Tensor, please make sure that the model can run with model((calib_data[0]))
2025-12-16 20:47:45,459 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,460 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,461 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,461 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,461 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,462 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,462 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,462 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,463 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,463 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,463 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,464 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,464 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,464 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,465 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,465 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,465 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,466 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,466 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,466 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,467 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,467 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,467 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,468 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,468 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,468 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,469 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,469 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,469 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,470 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,470 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,470 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLVisionBlocktoQuantQwen25VLVisionBlock2025-12-16 20:47:45,471 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,471 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,472 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,472 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,472 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,473 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,473 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,473 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,474 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,474 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,474 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,475 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,475 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,476 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,476 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,476 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,477 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,477 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,477 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,478 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,478 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,478 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,479 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,479 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,479 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,480 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,480 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer2025-12-16 20:47:45,480 - msmodelslim - INFO - multimodal model replace block
Qwen2_5_VLDecoderLayertoQuantQwen2VLDecoderLayer...[W compiler_depend.ts:82] Warning: The oprator of add is executed, Currently High Accuracy but Low Performance OP with 64-bit has been used, Please Do Some Cast at Python Functions with 32-bit for Better Performance! (function operator())
../usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py:518: UserWarning: current tensor is running as_strided, don't perform inplace operations on the returned value. If you encounter this warning and have precision issues, you can try torch.npu.config.allow_internal_format = False to resolve precision issues. (Triggered internally at build/CMakeFiles/torch_npu.dir/compiler_depend.ts:128.)
hidden_states = hidden_states[window_index, :, :]
.2025-12-16 20:48:44,938 - msmodelslim - INFO - current block is QuantQwen25VLVisionBlock, layername:visual.blocks.0
...........[W compiler_depend.ts:79] Warning: [Check][offset] Check input storage_offset[%ld] = 0 failed, result is untrustworthy59968 (function operator())
............Traceback (most recent call last):
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_outlier.py", line 409, in process
self._process()
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_outlier.py", line 530, in _process
self.model(*data)
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py", line 1757, in forward
image_embeds = self.visual(pixel_values, grid_thw=image_grid_thw)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py", line 546, in forward
hidden_states = blk(hidden_states, cu_seqlens=cu_seqlens_now, position_embeddings=position_embeddings)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_block.py", line 795, in forward
hidden_states = hidden_states + self.attn(
^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/transformers/models/qwen2_5_vl/modeling_qwen2_5_vl.py", line 285, in forward
q, k, v = self.qkv(hidden_states).reshape(seq_length, 3, self.num_heads, -1).permute(1, 0, 2, 3).unbind(0)
^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib64/python3.11/site-packages/torch/nn/modules/linear.py", line 114, in forward
return F.linear(input, self.weight, self.bias)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
RuntimeError: expected dtype Float but got dtype Half
[ERROR] 2025-12-16-20:52:38 (PID:10893, Device:0, RankID:-1) ERR01002 OPS invalid type
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/workspace/msit/msmodelslim/example/multimodal_vlm/Qwen2.5-VL/quant_qwen2_5vl.py", line 120, in
anti_outlier.process()
File "/usr/local/lib64/python3.11/site-packages/torch/utils/_contextlib.py", line 115, in decorate_context
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.11/site-packages/msmodelslim/pytorch/llm_ptq/anti_outlier/anti_outlier.py", line 411, in process
raise Exception("Please check your config, model and input!", e) from e
Exception: ('Please check your config, model and input!', RuntimeError('expected dtype Float but got dtype Half\n[ERROR] 2025-12-16-20:52:38 (PID:10893, Device:0, RankID:-1) ERR01002 OPS invalid type'))
Special notes for this issue/备注 (Optional / 选填)