MindSpeed-LLM 预置稠密大模型
认证【Pass】表示经过昇腾官方版本测试的模型。【Test】表示模型处于内部测试阶段,未完成充分的性能测试和验收,在实际使用中可能存在未被发现的问题,待后续充分验证后会发布正式版本。相关使用问题可反馈至MindSpeed-LLM/issues。
| 模型 | 下载链接 | 脚本位置 | 序列 | 实现 | 集群 | 贡献方 | 认证 |
|---|---|---|---|---|---|---|---|
| Aquila | 7B | aquila | 2K | Legacy | 1x8 | 【GTS】 | 【Pass】 |
| Aquila2 | 7B | aquila2 | 2K | Legacy | 1x8 | 【GTS】 | 【Pass】 |
| 34B | 4K | Legacy | 2x8 | 【GTS】 | 【Pass】 | ||
| Baichuan | 7B | baichuan | 4K | Legacy | 1x8 | 【GTS】 | 【Pass】 |
| 13B | 4K | Legacy | 1x8 | 【GTS】 | 【Pass】 | ||
| Baichuan2 | 7B | baichuan2 | 4K | Legacy | 1x8 | 【Ascend】 | 【Pass】 |
| 13B | 4K | Mcore | 1x8 | 【Ascend】 | 【Pass】 | ||
| Bloom | 7B1 | bloom | 2K | Legacy | 1x8 | 【Ascend】 | 【Pass】 |
| 176B | 2K | Legacy | 12x8 | 【Ascend】 | 【Pass】 | ||
| ChatGLM3 | 6B | chatglm3 | 8K | Mcore | 1x8 | 【Ascend】 | 【Pass】 |
| 32K | Mcore | 1x8 | 【Ascend】 | 【Pass】 | |||
| 64K | Mcore | 2x8 | 【Ascend】 | 【Pass】 | |||
| GLM4 | 9B | glm4 | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 32K | Mcore | 2x8 | 【GTS】 | 【Pass】 | |||
| CodeLlama | 34B | codellama | 4K | Mcore | 2x8 | 【GTS】 | 【Pass】 |
| InternLM | 7B | intern | 2K | Legacy | 1x8 | 【Ascend】 | 【Pass】 |
| 65B | 2K | Legacy | 4x8 | 【Ascend】 | 【Pass】 | ||
| InternLM2 | 20B | internlm2 | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | |||
| InternLM2.5 | 1.8B | internlm25 | 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 7B | 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 20B | 32K | Mcore | 2x8 | 【GTS】 | 【Test】 | ||
| LLaMA | 7B | llama | 2K | Legacy | 1x8 | 【Ascend】 | 【Pass】 |
| 13B | 2K | Legacy | 1x8 | 【Ascend】 | 【Pass】 | ||
| 33B | 2K | Legacy | 4x8 | 【Ascend】 | 【Pass】 | ||
| 65B | 2K | Legacy | 4x8 | 【Ascend】 | 【Pass】 | ||
| LLaMA2 | 7B | llama2 | 4K | Mcore | 1x8 | 【NAIE】 | 【Pass】 |
| 13B | 4K | Mcore | 1x8 | 【NAIE】 | 【Pass】 | ||
| 34B | 4K | Mcore | 2x8 | 【GTS】 | 【Pass】 | ||
| 70B | 4K | Mcore | 4x8 | 【GTS】 | 【Pass】 | ||
| LLaMA3 | 8B | llama3 | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 70B | 8K | Mcore | 4x8 | 【GTS】 | 【Pass】 | ||
| LLaMA3.1 | 8B | llama31 | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 128K | Mcore | 4x8 | 【GTS】 | 【Pass】 | |||
| 70B | 8K | Mcore | 4x8 | 【GTS】 | 【Pass】 | ||
| LLaMA3.2 | 1B | llama32 | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 3B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| LLaMA3.3 | 70B-Instruct | llama33 | 8K | Mcore | 4x8 | 【GTS】 | 【Test】 |
| Qwen | 7B | qwen | 8K | Legacy | 1x8 | 【GTS】 | 【Pass】 |
| 14B | 2K | Legacy | 1x8 | 【GTS】 | 【Pass】 | ||
| 72B | 8K | Legacy | 16x8 | 【GTS】 | 【Pass】 | ||
| Qwen1.5 | 0.5B | qwen15 | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 1.8B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 4B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 7B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 14B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 32B | 8K | Mcore | 4x8 | 【GTS】 | 【Pass】 | ||
| 72B | 8K | Mcore | 8x8 | 【GTS】 | 【Pass】 | ||
| 110B | 8K | Mcore | 8x8 | 【GTS】 | 【Pass】 | CodeQwen1.5 | 7B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| Qwen2 | 0.5B | qwen2 | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | |||
| 1.5B | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | |||
| 7B | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | |||
| 72B | 4K | Mcore | 4x8 | 【GTS】 | 【Pass】 | ||
| Qwen2.5 | 0.5B | qwen25 | 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | 1.5B | 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 | 3B | 32K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 7B | 32K | Mcore | 1x8 | 【Ascend】 | 【Pass】 | ||
| 14B | 32K | Mcore | 2x8 | 【GTS】 | 【Pass】 | ||
| 32B | 32K | Mcore | 4x8 | 【GTS】 | 【Pass】 | ||
| 72B | 32K | Mcore | 8x8 | 【GTS】 | 【Test】 | ||
| Qwen2.5-Math | 1.5B | qwen25_math | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 7B | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 72B | 4K | Mcore | 4x8 | 【GTS】 | 【Test】 | CodeQwen2.5 | 7B | qwen25_coder | 8K | Mcore | 1x8 | 【China Mobile Cloud】 | 【Test】 |
| Yi | 9B | yi | 4K | Legacy | 1x4 | 【OpenMind】 | 【Test】 |
| 34B | 4K | Mcore | 2x8 | 【GTS】 | 【Pass】 | ||
| Yi1.5 | 6B | yi15 | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 9B | 4K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| 34B | 4K | Mcore | 2x8 | 【GTS】 | 【Test】 | ||
| Mistral | 7B | mistral | 32K | Mcore | 1x8 | 【NAIE】 | 【Pass】 |
| Gemma | 2B | gemma | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 7B | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 | ||
| Gemma2 | 9B | gemma2 | 8K | Mcore | 1x8 | 【GTS】 | 【Pass】 |
| 27B | 8K | Mcore | 2x8 | 【GTS】 | 【Pass】 | ||
| GPT3 | 175B | gpt3 | 2K | Legacy | 16x8 | 【Ascend】 | 【Pass】 |
| MiniCPM | 2B | minicpm | 4K | Mcore | 1x8 | 【NAIE】 | 【Pass】 |
| MiniCPM3 | 4B | minicpm3 | 32K | Mcore | 1x8 | 【GTS】 | 【Test】 |
| Phi3.5 | mini-instruct | phi35 | 4K | Mcore | 1x8 | 【GTS】 | 【Test】 |
| DeepSeek-Math | 7B | deepseek_math | 4K | Mcore | 1x8 | 【Ascend】 | 【Test】 |
| DeepSeek-R1-Distill-Qwen | 1.5B | deepseek_r1_distill_qwen | 4K | Mcore | 1x8 | 【Ascend】 | 【Test】 |
| 7B | 4K | Mcore | 1x8 | 【Ascend】 | 【Test】 | ||
| 14B | 4K | Mcore | 1x8 | 【Ascend】 | 【Test】 | ||
| 32B | 8K | Mcore | 2x8 | 【Ascend】 | 【Test】 | ||
| DeepSeek-R1-Distill-LLaMA | 8B | deepseek_r1_distill_llama | 8K | Mcore | 1x8 | 【Ascend】 | 【Test】 |
| 70B | 8K | Mcore | 4x8 | 【Ascend】 | 【Test】 |
以上模型脚本环境变量声明:
ASCEND_LAUNCH_BLOCKING:将Host日志输出到串口,0-关闭/1-开启
ASCEND_SLOG_PRINT_TO_STDOUT:设置默认日志级别,0-debug/1-info/2-warning/3-error
HCCL_WHITELIST_DISABLE:HCCL白名单开关,1-关闭/0-开启
HCCL_CONNECT_TIMEOUT:设置HCCL超时时间,默认值为120
CUDA_DEVICE_MAX_CONNECTIONS:定义了任务流能够利用或映射到的硬件队列的数量
TASK_QUEUE_ENABLE: 用于控制开启task_queue算子下发队列优化的等级,0:关闭,1:开启Level 1优化,2:开启Level 2优化
COMBINED_ENABLE: 设置combined标志。设置为0表示关闭此功能;设置为1表示开启,用于优化非连续两个算子组合类场景
PYTORCH_NPU_ALLOC_CONF:内存碎片优化开关,默认是expandable_segments:False,使能时expandable_segments:True
ASCEND_RT_VISIBLE_DEVICES:指定哪些Device对当前进程可见,支持一次指定一个或多个Device ID。通过该环境变量,可实现不修改应用程序即可调整所用Device的功能。
NPUS_PER_NODE: 配置一个计算节点上使用的NPU数量
社区BUG列表
-
Baichuan-13B: 在任务执行过程中如果出现报错:AttributeError: 'BaichuanTokenizer’ object has no attribute 'sp_model',请执行下面命令解决这个问题:
pip install transformers==4.32.0 --force -
GPT: GPT词表文件与常规模型不同:
mkdir vocab_file cd vocab_file wget https://s3.amazonaws.com/models.huggingface.co/bert/gpt2-vocab.json wget https://s3.amazonaws.com/models.huggingface.co/bert/gpt2-merges.txt cd .. # 处理成训练数据 python ./preprocess_data.py \ --input ./dataset/ \ --output-prefix ./dataset/gpt_text_sentence \ --tokenizer-type GPT2BPETokenizer \ --vocab-file ./vocab_file/gpt2-vocab.json \ --merge-file ./vocab_file/gpt2-merges.txt \ --append-eod \ --workers 4 \ --log-interval 1000 # 请根据真实存放路径配置预训练脚本以下参数 VOCAB_FILE="./vocab_file/gpt2-vocab.json" # 词表 MERGE_FILE="./vocab_file/gpt2-merges.txt" # BPE 合并表 DATA_PATH="./dataset/gpt_text_sentence" # 数据路径 -
Bloom-176B: config.json中同字段对应的key值与其他模型不一致,将文件中的n_embed改为hidden_size, 将num_attention_heads修改为n_head
-
QWen: 不包含QWen1.5等,需要修改权重文件
# 修改modelling_qwen.py文件第39行,将: # SUPPORT_FP16 = SUPPORT_CUDA and torch.cuda.get_device_capability(0)[0] >= 7 # 修改为: # SUPPORT_FP16 = True
版本支持
1.Llama3.3系列模型依赖以下特定Python及transformers版本
Python>=3.10
transformers==4.51.0