已合并
add base model: gliner_large-v2.5 #983
lilinjie11创建于 6月25日
add base model: gliner_large-v2.5 #983
已合并
共 8 个文件变更+1076-25
| @@ -260,6 +260,7 @@ ge.exec.precision_mode=force_fp32 | |||
| 260 | 3. **图像预处理使用 `AutoImageProcessor(..., use_fast=False)`** — 慢版 processor 支持 `return_tensors="np"` | 260 | 3. **图像预处理使用 `AutoImageProcessor(..., use_fast=False)`** — 慢版 processor 支持 `return_tensors="np"` |
| 261 | 4. **多模态图像 token 用字符串构建** — 如 `<|vision_start|><|image_pad|>...<|vision_end|>`,不依赖 processor 的图像处理 | 261 | 4. **多模态图像 token 用字符串构建** — 如 `<|vision_start|><|image_pad|>...<|vision_end|>`,不依赖 processor 的图像处理 |
| 262 | 5. **不在代码中指定 `context.ascend.precision_mode`** — 精度模式由 `converter_lite` 转换时的 `config.ini`(如 `force_fp32`)控制,推理脚本中不应重复设置,避免与转换配置冲突 | 262 | 5. **不在代码中指定 `context.ascend.precision_mode`** — 精度模式由 `converter_lite` 转换时的 `config.ini`(如 `force_fp32`)控制,推理脚本中不应重复设置,避免与转换配置冲突 |
| 263 | +6. **`Model.predict` 返回值直接取用,无需类型校验** — `mslite.Model.predict()` 返回值一定是 `MSTensor` 列表,每个元素必然具备 `get_data_to_numpy()` 方法。直接调用 `outputs[0].get_data_to_numpy()` 取 numpy 数组即可,不要写 `hasattr(outputs[0], "get_data_to_numpy")` 之类的分支判断,也不要保留 `np.array(outputs[0])` 回退——MindSpore Lite 接口不会返回 numpy 类型,此类校验属于无效代码 | ||
| 263 | 264 | ||
| 264 | ### 5.2 输入 dtype 对齐 | 265 | ### 5.2 输入 dtype 对齐 |
| 265 | 266 | ||
M.claude/skills/lite-cloud-side/open-source-model-migration/references/update_model_support_table.md+61-7
| @@ -30,9 +30,17 @@ description: Update MindSpore Lite README_CN.md/README.md supported model tables | |||
| 30 | 30 | ||
| 31 | 1. 盘点模型目录 | 31 | 1. 盘点模型目录 |
| 32 | - 列出 `mindspore-lite/examples/base_models/` 下**每个子目录**(必要时包含二级子目录作为具体模型,例如 `yolov10/yolov10-X`)。 | 32 | - 列出 `mindspore-lite/examples/base_models/` 下**每个子目录**(必要时包含二级子目录作为具体模型,例如 `yolov10/yolov10-X`)。 |
| 33 | + - 推荐命令(自动排除 `configs`、`utils`、`upstream` 等组织性子目录): | ||
| 34 | + ```bash | ||
| 35 | + find mindspore-lite/examples/base_models \ | ||
| 36 | + -mindepth 1 -maxdepth 2 -type d \ | ||
| 37 | + -not -name configs -not -name utils -not -name upstream \ | ||
| 38 | + | sort | ||
| 39 | + ``` | ||
| 40 | + - 跳过**空目录/占位目录**:若子目录内既无 `README*` 也无 `*.py` / `*.onnx` / 权重文件,视为未落地,不计入表格(如 `tcp/upstream/TCP/` 这种只有空嵌套的)。 | ||
| 33 | 2. 对照表格现有条目 | 41 | 2. 对照表格现有条目 |
| 34 | - 在两份 README 的表格中搜索每个模型: | 42 | - 在两份 README 的表格中搜索每个模型: |
| 35 | - - 找到:仅做“补链接 + `✅`”(不改模型名文本,除非表格明显不一致) | 43 | + - 找到:仅做“补链接 + `✅`”(不改模型名文本,除非触发下方“明显不一致判定”) |
| 36 | - 找不到:进入“新增条目”流程 | 44 | - 找不到:进入“新增条目”流程 |
| 37 | 3. 生成链接 | 45 | 3. 生成链接 |
| 38 | - 统一使用 AtomGit tree 链接,路径为 base_models 下相对路径: | 46 | - 统一使用 AtomGit tree 链接,路径为 base_models 下相对路径: |
| @@ -44,14 +52,28 @@ description: Update MindSpore Lite README_CN.md/README.md supported model tables | |||
| 44 | - 只填目标列这一格,其它列保持原值不动。 | 52 | - 只填目标列这一格,其它列保持原值不动。 |
| 45 | - 当目标列所有行都非空时,才在表格末尾新增一行: | 53 | - 当目标列所有行都非空时,才在表格末尾新增一行: |
| 46 | - 新行其它列留空,目标列写入新模型的“链接 + `✅`”。 | 54 | - 新行其它列留空,目标列写入新模型的“链接 + `✅`”。 |
| 55 | + - 一次新增多个条目到**同一列**时,按“**子类聚类 + 字母序**”决定填入空单元格的先后: | ||
| 56 | + - 例:同列已有 `yolov10x`、`vit-*`、`bert-*`,新增 `yolov8`、`gliner-*`、`grounding-dino-*` 时,优先把 `yolov8` 放到离 `yolov10x` 最近的空单元格;其余按目录名字母序自上而下填充。 | ||
| 57 | + - 找不到子类聚类关系时,统一按目录名字母序填充。 | ||
| 58 | + - 不必为"聚类"强行插入新行或挪动已有条目。 | ||
| 47 | 5. 中英文同步 | 59 | 5. 中英文同步 |
| 48 | - 在 `README_CN.md` 完成更新后,把**同样的结构变更**同步到 `README.md`: | 60 | - 在 `README_CN.md` 完成更新后,把**同样的结构变更**同步到 `README.md`: |
| 49 | - 相同模型的链接必须一致 | 61 | - 相同模型的链接必须一致 |
| 50 | - 新增规则一致(同一列的空单元格复用逻辑一致) | 62 | - 新增规则一致(同一列的空单元格复用逻辑一致) |
| 51 | 6. 自检 | 63 | 6. 自检 |
| 52 | - - 每一行的 `|` 列数一致(5 列) | 64 | + - 每一行的 `|` 列数一致(**6 列**)。一行 awk 即可核对(无输出即正确): |
| 53 | - - 新增内容没有引入多余的空白行或多余的表格行 | 65 | + ```bash |
| 54 | - - 链接路径与目录实际存在的相对路径一致 | 66 | + awk -F'|' '/云侧推理模型支持列表/,/^### API与文档/' README_CN.md \ |
| 67 | + | awk -F'|' 'NF>1 && NF-2!=6 {print NR": "NF-2" cols — "$0}' | ||
| 68 | + ``` | ||
| 69 | + EN 版把章节标题换成 `Supported models for cloud-side` / `^### API and documentation`。 | ||
| 70 | + - 新增内容没有引入多余的空白行或多余的表格行。 | ||
| 71 | + - 链接路径与目录实际存在的相对路径一致。批量校验所有链接: | ||
| 72 | + ```bash | ||
| 73 | + grep -oE 'tree/master/mindspore-lite/examples/base_models/[^)]+' README_CN.md README.md \ | ||
| 74 | + | sed 's|.*tree/master/||' | sort -u \ | ||
| 75 | + | while read -r p; do [ -e "$p" ] || echo "MISSING: $p"; done | ||
| 76 | + ``` | ||
| 55 | 77 | ||
| 56 | ## 列分类规则(可扩展) | 78 | ## 列分类规则(可扩展) |
| 57 | 79 | ||
| @@ -62,8 +84,11 @@ description: Update MindSpore Lite README_CN.md/README.md supported model tables | |||
| 62 | - 信息检索/向量嵌入/CNN/其他 | 84 | - 信息检索/向量嵌入/CNN/其他 |
| 63 | - 目录名/模型名含:`reranker`、`embedding`、`vit`、`yolo`、`bevdet`、`bert` | 85 | - 目录名/模型名含:`reranker`、`embedding`、`vit`、`yolo`、`bevdet`、`bert` |
| 64 | - 视觉语言模型(VLM) | 86 | - 视觉语言模型(VLM) |
| 65 | - - 模型名含 `VL` 且语义为 VLM(如 `...VL...Instruct`) | 87 | + - 满足**任一**即归入 VLM: |
| 66 | - - 注意:如果表格历史上把某些 `VL-Embedding` / `VL-Reranker` 放在“其他”,则新增时保持一致放“其他” | 88 | + - 模型名含 `VL` 且语义为 VLM(如 `...VL...Instruct`); |
| 89 | + - 任务语义为视觉-语言:OCR(`ocr`)、文本条件检测/grounding(`grounding_dino`)、image-caption、VQA 等。 | ||
| 90 | + - 注意:如果表格历史上把某些 `VL-Embedding` / `VL-Reranker` 放在“其他”,则新增时保持一致放“其他”。 | ||
| 91 | + - VLM 列已满(需要新增行)而“其他”列仍有空单元格时,可酌情把**非对话型 VLM**(OCR、grounding、CLIP 类视觉编码器)放到“其他”,避免新增行;对话型 VLM(`*-Instruct` / `*-Thinking`)仍优先放 VLM 列。 | ||
| 67 | - 大语言模型(LLM) | 92 | - 大语言模型(LLM) |
| 68 | - `qwen` 系列且不属于 `vl/asr/tts/reranker/embedding` | 93 | - `qwen` 系列且不属于 `vl/asr/tts/reranker/embedding` |
| 69 | - 图像/视频生成模型 | 94 | - 图像/视频生成模型 |
| @@ -71,7 +96,15 @@ description: Update MindSpore Lite README_CN.md/README.md supported model tables | |||
| 71 | 96 | ||
| 72 | ## 常见目录名 ↔ 展示名(建议) | 97 | ## 常见目录名 ↔ 展示名(建议) |
| 73 | 98 | ||
| 74 | -新增条目时,若表格中没有该模型名,可参考以下转换生成展示名(也可直接用目录名作为展示名): | 99 | +新增条目时,若表格中没有该模型名,按以下**主规则**生成展示名: |
| 100 | + | ||
| 101 | +- **主规则(默认)**:沿用目录名,仅把 `_` 换成 `-`,保持原大小写。 | ||
| 102 | + - 例:`bert_base_chinese` → `bert-base-chinese`、`yolov8` → `yolov8`、`gliner_large-v2.5` → `gliner-large-v2.5`、`glm_ocr` → `glm-ocr`。 | ||
| 103 | +- **例外 1(有广为人知的官方品牌大小写时启用)**:`ViT-*`、`BEVDet`、`YOLOv*`、`GLiNER-*`、`GLM-*`、`Grounding-DINO-*` 等可改用官方大小写。 | ||
| 104 | +- **例外 2(已有同族模型时跟随)**:表格里已存在 `yolov10x`(小写)→ 新增 `yolov8` 也用小写,不要写成 `YOLOv8`;已存在 `Qwen3-*`(Title Case)→ 新增同族 qwen 模型也用 Title Case。 | ||
| 105 | +- 拿不准时回退主规则(沿用目录名最安全)。 | ||
| 106 | + | ||
| 107 | +历史样例(仅供回溯参考,新条目以上述主规则为准): | ||
| 75 | 108 | ||
| 76 | - `qwen3.5_4b` → `Qwen3.5-4B` | 109 | - `qwen3.5_4b` → `Qwen3.5-4B` |
| 77 | - `qwen3_5_0.8b` → `Qwen3.5-0.8B` | 110 | - `qwen3_5_0.8b` → `Qwen3.5-0.8B` |
| @@ -88,6 +121,27 @@ description: Update MindSpore Lite README_CN.md/README.md supported model tables | |||
| 88 | - `bert_base_chinese` → `bert_base_chinese` | 121 | - `bert_base_chinese` → `bert_base_chinese` |
| 89 | - `yolov10/yolov10-X` → `yolov10-X` | 122 | - `yolov10/yolov10-X` → `yolov10-X` |
| 90 | 123 | ||
| 124 | +## 明显不一致判定(决定"是否改显示文本") | ||
| 125 | + | ||
| 126 | +仅当满足以下**任一**才视为"明显不一致",可改表格显示文本;否则一律保留原文本只补链接: | ||
| 127 | + | ||
| 128 | +- 显示名中的**关键修饰词**在目录名里完全找不到。例:显示 `Qwen3-VL-Reranker-8B`、目录却是 `qwen3_reranker_8b`(无 `vl`)→ 可考虑去掉 `VL`。 | ||
| 129 | +- 显示名与目录名指向**不同尺寸/版本**。例:显示 `Qwen3-VL-4B-Instruct`、目录却是 `qwen3_vl_4b_thinking`。 | ||
| 130 | +- 显示名有明显拼写错误。例:`Kand0-T2V0` 这类历史笔误(属"约定俗成"的可不动)。 | ||
| 131 | + | ||
| 132 | +边界模糊时**保守处理**:只补链接,并在交付总结里单列疑点让用户复核,不擅自改文本。 | ||
| 133 | + | ||
| 134 | +## 反向不一致(表格条目无对应目录) | ||
| 135 | + | ||
| 136 | +同步是单向"目录 → 表格"。遇到反向不一致时: | ||
| 137 | + | ||
| 138 | +1. **表格有条目、目录里没有**: | ||
| 139 | + - **不要自行删除**(可能在其他仓库维护、或尚未开源、或表格为路线图)。 | ||
| 140 | + - 保留原条目不动,**不补** `✅`。 | ||
| 141 | + - 在交付总结里单列"无目录的表格条目",让用户决定下架/保留。 | ||
| 142 | +2. **表格有条目、目录对得上但有疑点**(如显示 `Qwen3-VL-Reranker-8B`、目录是 `qwen3_reranker_8b`): | ||
| 143 | + - 按"明显不一致判定"处理;判定不通过则只补链接、不改文本,并单列疑点。 | ||
| 144 | + | ||
| 91 | ## 单元格写法规范 | 145 | ## 单元格写法规范 |
| 92 | 146 | ||
| 93 | - 已存在文本补链接:`[展示名](URL) ✅` | 147 | - 已存在文本补链接:`[展示名](URL) ✅` |
| @@ -89,24 +89,24 @@ If you wish to further learn and use MindSpore Lite, please refer to the followi | |||
| 89 | | Image/Video Generation Models | Vision-Language Models (VLM) | Large Language Models (LLM) | Audio Models (ASR/TTS) | Autonomous Driving / Embodied Intelligence | Information Retrieval / Embeddings / CNN / Others | | 89 | | Image/Video Generation Models | Vision-Language Models (VLM) | Large Language Models (LLM) | Audio Models (ASR/TTS) | Autonomous Driving / Embodied Intelligence | Information Retrieval / Embeddings / CNN / Others | |
| 90 | | :---------------------------: | :--------------------------: | :-------------------------: | :-------------------: | :----------------------------------------: | :----------------------------------------------: | | 90 | | :---------------------------: | :--------------------------: | :-------------------------: | :-------------------: | :----------------------------------------: | :----------------------------------------------: | |
| 91 | | Kandinsky-5.0-I2V-Lite-5s | Qwen3-VL-8B-Thinking | Qwen3.6-27B | WeNet | Mask2Former | [Qwen3-Reranker-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_4b) ✅ | | 91 | | Kandinsky-5.0-I2V-Lite-5s | Qwen3-VL-8B-Thinking | Qwen3.6-27B | WeNet | Mask2Former | [Qwen3-Reranker-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_4b) ✅ | |
| 92 | -| Kand0-T2V0-T2V-Lite-sft-10s | Qwen3-VL-8B-Instruct | Qwen3.5-27B | [FireRedASR-AED-L](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/fireredasr_aed_l) ✅ | DinoV3 | Qwen3-VL-Reranker-8B | | 92 | +| Kand0-T2V0-T2V-Lite-sft-10s | [Qwen3-VL-8B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_8b_instruct) ✅ | Qwen3.5-27B | [FireRedASR-AED-L](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/fireredasr_aed_l) ✅ | DinoV3 | [Qwen3-VL-Reranker-8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_8b) ✅ | |
| 93 | | Kandinsky-5.0-T2I-Lite | [Qwen3-VL-4B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_thinking) ✅ | Qwen3.5-9B | [CosyVoice2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/cosyvoice2_0.5b) ✅ | CenterPoint(2D) | [Qwen3-VL-Reranker-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_reranker_2b) ✅ | | 93 | | Kandinsky-5.0-T2I-Lite | [Qwen3-VL-4B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_thinking) ✅ | Qwen3.5-9B | [CosyVoice2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/cosyvoice2_0.5b) ✅ | CenterPoint(2D) | [Qwen3-VL-Reranker-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_reranker_2b) ✅ | |
| 94 | | Kandinsky-5.0-I2I-Lite | [Qwen3-VL-4B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_instruct) ✅ | [Qwen3.5-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_4b) ✅ | CosyVoice3-0.5B | CenterPoint(3D) | Qwen3-VL-Embedding-8B | | 94 | | Kandinsky-5.0-I2I-Lite | [Qwen3-VL-4B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_instruct) ✅ | [Qwen3.5-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_4b) ✅ | CosyVoice3-0.5B | CenterPoint(3D) | Qwen3-VL-Embedding-8B | |
| 95 | | Wan2.1-T2V-1.3B | [Qwen3-VL-2B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_thinking) ✅ | [Qwen3.5-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_2b) ✅ | Qwen3-ASR-0.6B | [bevdet-r50](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bevdet) ✅ | [Qwen3-VL-Embedding-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_embedding_2b) ✅ | | 95 | | Wan2.1-T2V-1.3B | [Qwen3-VL-2B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_thinking) ✅ | [Qwen3.5-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_2b) ✅ | Qwen3-ASR-0.6B | [bevdet-r50](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bevdet) ✅ | [Qwen3-VL-Embedding-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_embedding_2b) ✅ | |
| 96 | | Wan2.1-T2V-14B | [Qwen3-VL-2B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_instruct) ✅ | [Qwen3.5-0.8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_5_0.8b) ✅ | [Qwen3-ASR-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_asr_1.7b) ✅ | flashOCC | [Qwen3-Reranker-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_0.6b) ✅ | | 96 | | Wan2.1-T2V-14B | [Qwen3-VL-2B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_instruct) ✅ | [Qwen3.5-0.8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_5_0.8b) ✅ | [Qwen3-ASR-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_asr_1.7b) ✅ | flashOCC | [Qwen3-Reranker-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_0.6b) ✅ | |
| 97 | | Wan2.1-I2V-14B-480P | [Qwen2.5-VL-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_vl_3b_instruct) ✅ | Qwen3-30B-A3B | Qwen3-TTS-12Hz-1.7B-Base | TinyVLA | [jina-reranker-v3](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/jina_reranker_v3) ✅ | | 97 | | Wan2.1-I2V-14B-480P | [Qwen2.5-VL-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_vl_3b_instruct) ✅ | Qwen3-30B-A3B | Qwen3-TTS-12Hz-1.7B-Base | TinyVLA | [jina-reranker-v3](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/jina_reranker_v3) ✅ | |
| 98 | -| Wan2.2-TI2V-5B | Qwen2-VL-2B-Instruct | Qwen3-8B | [Qwen3-TTS-12Hz-1.7B-CustomVoice](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_tts_12hz_1.7b_customvoice) ✅ | GR00TN1.7 | [Qwen3-Embedding-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_0_6b) ✅ | | 98 | +| Wan2.2-TI2V-5B | Qwen2-VL-2B-Instruct | [Qwen3-8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_8b) ✅ | [Qwen3-TTS-12Hz-1.7B-CustomVoice](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_tts_12hz_1.7b_customvoice) ✅ | GR00TN1.7 | [Qwen3-Embedding-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_0_6b) ✅ | |
| 99 | | Wan2.2-T2V-A14B | Qwen2-VL-2B | [Qwen3-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_4b) ✅ | Qwen3-TTS-12Hz-1.7B-VoiceDesign | SpatialVLA | [Qwen3-Embedding-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_4b) ✅ | | 99 | | Wan2.2-T2V-A14B | Qwen2-VL-2B | [Qwen3-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_4b) ✅ | Qwen3-TTS-12Hz-1.7B-VoiceDesign | SpatialVLA | [Qwen3-Embedding-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_4b) ✅ | |
| 100 | | Wan2.2-I2V-A14B | InternVL3_5-4B-Flash | [Qwen3-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_1.7b) ✅ | Qwen3-TTS-12Hz-0.6B-Base | SmolVLA | [yolov10x](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/yolov10/yolov10-X) ✅ | | 100 | | Wan2.2-I2V-A14B | InternVL3_5-4B-Flash | [Qwen3-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_1.7b) ✅ | Qwen3-TTS-12Hz-0.6B-Base | SmolVLA | [yolov10x](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/yolov10/yolov10-X) ✅ | |
| 101 | | Wan2.2-Animate-14B | InternVL3_5-2B-Flash | [Qwen3-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_0.6b) ✅ | Qwen3-TTS-12Hz-0.6B-CustomVoice | MiniVLA | [vit-base-patch16-224](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/vit_base_patch16_224) ✅ | | 101 | | Wan2.2-Animate-14B | InternVL3_5-2B-Flash | [Qwen3-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_0.6b) ✅ | Qwen3-TTS-12Hz-0.6B-CustomVoice | MiniVLA | [vit-base-patch16-224](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/vit_base_patch16_224) ✅ | |
| 102 | | Qwen-Image-Edit | InternVL3_5-1B-Flash | [Qwen2.5-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_7b) ✅ | | openVLA | [bert-base-chinese](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bert_base_chinese) ✅ | | 102 | | Qwen-Image-Edit | InternVL3_5-1B-Flash | [Qwen2.5-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_7b) ✅ | | openVLA | [bert-base-chinese](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bert_base_chinese) ✅ | |
| 103 | -| Qwen-Image | InternVL3-2B | [Qwen2.5-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_3b) ✅ | | openpi pi 0.5 | [Qwen2.5-Math-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b) ✅ | | 103 | +| Qwen-Image | InternVL3-2B | [Qwen2.5-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_3b) ✅ | | [openpi pi 0.5](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/open_pi_0.5) ✅ | [Qwen2.5-Math-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b) ✅ | |
| 104 | | FLUX.1-dev | InternVL3-1B | [Qwen2.5-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_0.5b) ✅ | | | [Qwen2.5-Math-1.5B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b_instruct) ✅ | | 104 | | FLUX.1-dev | InternVL3-1B | [Qwen2.5-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_0.5b) ✅ | | | [Qwen2.5-Math-1.5B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b_instruct) ✅ | |
| 105 | | stable-diffusion-v1-5 | llava-v1.6 | [Qwen2-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_7b) ✅ | | | [Qwen3Guard-Gen-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3guard_gen_0_6b) ✅ | | 105 | | stable-diffusion-v1-5 | llava-v1.6 | [Qwen2-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_7b) ✅ | | | [Qwen3Guard-Gen-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3guard_gen_0_6b) ✅ | |
| 106 | -| stable-diffusion-2-1 | LLaVa | [Qwen2-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_1.5b) ✅ | | | | | 106 | +| stable-diffusion-2-1 | LLaVa | [Qwen2-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_1.5b) ✅ | | | [GLiNER-Large-v2.5](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/gliner_large-v2.5) ✅ | |
| 107 | -| stable-diffusion-xl-base-1.0 | BLIP | [Qwen2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_0.5b) ✅ | | | | | 107 | +| stable-diffusion-xl-base-1.0 | BLIP | [Qwen2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_0.5b) ✅ | | | [GLM-OCR](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/glm_ocr) ✅ | |
| 108 | -| | BLIP-2 | Qwen1.5-moe-a2.7B | | | | | 108 | +| | BLIP-2 | Qwen1.5-moe-a2.7B | | | [Grounding-DINO-Base](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/grounding_dino_base) ✅ | |
| 109 | -| | CLIP | | | | | | 109 | +| | [CLIP](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/clip_vit_base_patch32) ✅ | | | | [yolov8](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/yolov8) ✅ | |
| 110 | 110 | ||
| 111 | ### API and documentation | 111 | ### API and documentation |
| 112 | 112 | ||
| @@ -148,8 +148,8 @@ If you wish to further learn and use MindSpore Lite, please refer to the followi | |||
| 148 | 148 | ||
| 149 | - [MindSpore](https://atomgit.com/mindspore/mindspore) | 149 | - [MindSpore](https://atomgit.com/mindspore/mindspore) |
| 150 | 150 | ||
| 151 | -- [MindOne](https://github.com/mindspore-lab/mindone) | 151 | +- [MindOne](https://atomgit.com/mindspore/mindone) |
| 152 | 152 | ||
| 153 | -- [Mindyolo](https://github.com/mindspore-lab/mindyolo) | 153 | +- [Mindyolo](https://atomgit.com/mindspore/mindyolo) |
| 154 | 154 | ||
| 155 | - [OpenHarmony](https://atomgit.com/openharmony/third_party_mindspore) | 155 | - [OpenHarmony](https://atomgit.com/openharmony/third_party_mindspore) |
| @@ -90,24 +90,24 @@ MindSpore Lite针对AIGC、语音类算法以及CV类模型推理,实现推理 | |||
| 90 | | 图像/视频生成模型 | 视觉语言模型(VLM) | 大语言模型(LLM) | 音频模型(ASR/TTS) | 自动驾驶/具身智能 | 信息检索/向量嵌入/CNN/其他模型 | | 90 | | 图像/视频生成模型 | 视觉语言模型(VLM) | 大语言模型(LLM) | 音频模型(ASR/TTS) | 自动驾驶/具身智能 | 信息检索/向量嵌入/CNN/其他模型 | |
| 91 | | :--------------: | :--------------: | :-------------: | :--------------: | :--------------: | :---------------------------: | | 91 | | :--------------: | :--------------: | :-------------: | :--------------: | :--------------: | :---------------------------: | |
| 92 | | Kandinsky-5.0-I2V-Lite-5s | Qwen3-VL-8B-Thinking | Qwen3.6-27B | WeNet | Mask2Former | [Qwen3-Reranker-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_4b) ✅ | | 92 | | Kandinsky-5.0-I2V-Lite-5s | Qwen3-VL-8B-Thinking | Qwen3.6-27B | WeNet | Mask2Former | [Qwen3-Reranker-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_4b) ✅ | |
| 93 | -| Kand0-T2V0-T2V-Lite-sft-10s | Qwen3-VL-8B-Instruct | Qwen3.5-27B | [FireRedASR-AED-L](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/fireredasr_aed_l) ✅ | DinoV3 | Qwen3-VL-Reranker-8B | | 93 | +| Kand0-T2V0-T2V-Lite-sft-10s | [Qwen3-VL-8B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_8b_instruct) ✅ | Qwen3.5-27B | [FireRedASR-AED-L](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/fireredasr_aed_l) ✅ | DinoV3 | [Qwen3-VL-Reranker-8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_8b) ✅ | |
| 94 | | Kandinsky-5.0-T2I-Lite | [Qwen3-VL-4B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_thinking) ✅ | Qwen3.5-9B | [CosyVoice2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/cosyvoice2_0.5b) ✅ | CenterPoint(2D) | [Qwen3-VL-Reranker-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_reranker_2b) ✅ | | 94 | | Kandinsky-5.0-T2I-Lite | [Qwen3-VL-4B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_thinking) ✅ | Qwen3.5-9B | [CosyVoice2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/cosyvoice2_0.5b) ✅ | CenterPoint(2D) | [Qwen3-VL-Reranker-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_reranker_2b) ✅ | |
| 95 | | Kandinsky-5.0-I2I-Lite | [Qwen3-VL-4B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_instruct) ✅ | [Qwen3.5-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_4b) ✅ | CosyVoice3-0.5B | CenterPoint(3D) | Qwen3-VL-Embedding-8B | | 95 | | Kandinsky-5.0-I2I-Lite | [Qwen3-VL-4B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_4b_instruct) ✅ | [Qwen3.5-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_4b) ✅ | CosyVoice3-0.5B | CenterPoint(3D) | Qwen3-VL-Embedding-8B | |
| 96 | | Wan2.1-T2V-1.3B | [Qwen3-VL-2B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_thinking) ✅ | [Qwen3.5-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_2b) ✅ | Qwen3-ASR-0.6B | [bevdet-r50](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bevdet) ✅ | [Qwen3-VL-Embedding-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_embedding_2b) ✅ | | 96 | | Wan2.1-T2V-1.3B | [Qwen3-VL-2B-Thinking](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_thinking) ✅ | [Qwen3.5-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3.5_2b) ✅ | Qwen3-ASR-0.6B | [bevdet-r50](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bevdet) ✅ | [Qwen3-VL-Embedding-2B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_embedding_2b) ✅ | |
| 97 | | Wan2.1-T2V-14B | [Qwen3-VL-2B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_instruct) ✅ | [Qwen3.5-0.8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_5_0.8b) ✅ | [Qwen3-ASR-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_asr_1.7b) ✅ | flashOCC | [Qwen3-Reranker-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_0.6b) ✅ | | 97 | | Wan2.1-T2V-14B | [Qwen3-VL-2B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_vl_2b_instruct) ✅ | [Qwen3.5-0.8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_5_0.8b) ✅ | [Qwen3-ASR-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_asr_1.7b) ✅ | flashOCC | [Qwen3-Reranker-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_reranker_0.6b) ✅ | |
| 98 | | Wan2.1-I2V-14B-480P | [Qwen2.5-VL-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_vl_3b_instruct) ✅ | Qwen3-30B-A3B | Qwen3-TTS-12Hz-1.7B-Base | TinyVLA | [jina-reranker-v3](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/jina_reranker_v3) ✅ | | 98 | | Wan2.1-I2V-14B-480P | [Qwen2.5-VL-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_vl_3b_instruct) ✅ | Qwen3-30B-A3B | Qwen3-TTS-12Hz-1.7B-Base | TinyVLA | [jina-reranker-v3](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/jina_reranker_v3) ✅ | |
| 99 | -| Wan2.2-TI2V-5B | Qwen2-VL-2B-Instruct | Qwen3-8B | [Qwen3-TTS-12Hz-1.7B-CustomVoice](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_tts_12hz_1.7b_customvoice) ✅ | GR00TN1.7 | [Qwen3-Embedding-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_0_6b) ✅ | | 99 | +| Wan2.2-TI2V-5B | Qwen2-VL-2B-Instruct | [Qwen3-8B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_8b) ✅ | [Qwen3-TTS-12Hz-1.7B-CustomVoice](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_tts_12hz_1.7b_customvoice) ✅ | GR00TN1.7 | [Qwen3-Embedding-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_0_6b) ✅ | |
| 100 | | Wan2.2-T2V-A14B | Qwen2-VL-2B | [Qwen3-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_4b) ✅ | Qwen3-TTS-12Hz-1.7B-VoiceDesign | SpatialVLA | [Qwen3-Embedding-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_4b) ✅ | | 100 | | Wan2.2-T2V-A14B | Qwen2-VL-2B | [Qwen3-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_4b) ✅ | Qwen3-TTS-12Hz-1.7B-VoiceDesign | SpatialVLA | [Qwen3-Embedding-4B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_embedding_4b) ✅ | |
| 101 | | Wan2.2-I2V-A14B | InternVL3_5-4B-Flash | [Qwen3-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_1.7b) ✅ | Qwen3-TTS-12Hz-0.6B-Base | SmolVLA | [yolov10x](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/yolov10/yolov10-X) ✅ | | 101 | | Wan2.2-I2V-A14B | InternVL3_5-4B-Flash | [Qwen3-1.7B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_1.7b) ✅ | Qwen3-TTS-12Hz-0.6B-Base | SmolVLA | [yolov10x](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/yolov10/yolov10-X) ✅ | |
| 102 | | Wan2.2-Animate-14B | InternVL3_5-2B-Flash | [Qwen3-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_0.6b) ✅ | Qwen3-TTS-12Hz-0.6B-CustomVoice | MiniVLA | [vit-base-patch16-224](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/vit_base_patch16_224) ✅ | | 102 | | Wan2.2-Animate-14B | InternVL3_5-2B-Flash | [Qwen3-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3_0.6b) ✅ | Qwen3-TTS-12Hz-0.6B-CustomVoice | MiniVLA | [vit-base-patch16-224](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/vit_base_patch16_224) ✅ | |
| 103 | | Qwen-Image-Edit | InternVL3_5-1B-Flash | [Qwen2.5-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_7b) ✅ | | openVLA | [bert-base-chinese](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bert_base_chinese) ✅ | | 103 | | Qwen-Image-Edit | InternVL3_5-1B-Flash | [Qwen2.5-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_7b) ✅ | | openVLA | [bert-base-chinese](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/bert_base_chinese) ✅ | |
| 104 | -| Qwen-Image | InternVL3-2B | [Qwen2.5-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_3b) ✅ | | openpi pi 0.5 | [Qwen2.5-Math-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b) ✅ | | 104 | +| Qwen-Image | InternVL3-2B | [Qwen2.5-3B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_3b) ✅ | | [openpi pi 0.5](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/open_pi_0.5) ✅ | [Qwen2.5-Math-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b) ✅ | |
| 105 | | FLUX.1-dev | InternVL3-1B | [Qwen2.5-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_0.5b) ✅ | | | [Qwen2.5-Math-1.5B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b_instruct) ✅ | | 105 | | FLUX.1-dev | InternVL3-1B | [Qwen2.5-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_0.5b) ✅ | | | [Qwen2.5-Math-1.5B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2.5_math_1.5b_instruct) ✅ | |
| 106 | | stable-diffusion-v1-5 | llava-v1.6 | [Qwen2-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_7b) ✅ | | | [Qwen3Guard-Gen-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3guard_gen_0_6b) ✅ | | 106 | | stable-diffusion-v1-5 | llava-v1.6 | [Qwen2-7B-Instruct](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_7b) ✅ | | | [Qwen3Guard-Gen-0.6B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen3guard_gen_0_6b) ✅ | |
| 107 | -| stable-diffusion-2-1 | LLaVa | [Qwen2-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_1.5b) ✅ | | | | | 107 | +| stable-diffusion-2-1 | LLaVa | [Qwen2-1.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_1.5b) ✅ | | | [GLiNER-Large-v2.5](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/gliner_large-v2.5) ✅ | |
| 108 | -| stable-diffusion-xl-base-1.0 | BLIP | [Qwen2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_0.5b) ✅ | | | | | 108 | +| stable-diffusion-xl-base-1.0 | BLIP | [Qwen2-0.5B](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/qwen2_0.5b) ✅ | | | [GLM-OCR](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/glm_ocr) ✅ | |
| 109 | -| | BLIP-2 | Qwen1.5-moe-a2.7B | | | | | 109 | +| | BLIP-2 | Qwen1.5-moe-a2.7B | | | [Grounding-DINO-Base](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/grounding_dino_base) ✅ | |
| 110 | -| | CLIP | | | | | | 110 | +| | [CLIP](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/clip_vit_base_patch32) ✅ | | | | [yolov8](https://atomgit.com/mindspore/mindspore-lite/tree/master/mindspore-lite/examples/base_models/yolov8) ✅ | |
| 111 | 111 | ||
| 112 | ### API与文档 | 112 | ### API与文档 |
| 113 | 113 | ||
| @@ -149,8 +149,8 @@ MindSpore Lite针对AIGC、语音类算法以及CV类模型推理,实现推理 | |||
| 149 | 149 | ||
| 150 | - [MindSpore](https://atomgit.com/mindspore/mindspore) | 150 | - [MindSpore](https://atomgit.com/mindspore/mindspore) |
| 151 | 151 | ||
| 152 | -- [MindOne](https://github.com/mindspore-lab/mindone) | 152 | +- [MindOne](https://atomgit.com/mindspore/mindone) |
| 153 | 153 | ||
| 154 | -- [Mindyolo](https://github.com/mindspore-lab/mindyolo) | 154 | +- [Mindyolo](https://atomgit.com/mindspore/mindyolo) |
| 155 | 155 | ||
| 156 | - [OpenHarmony](https://atomgit.com/openharmony/third_party_mindspore) | 156 | - [OpenHarmony](https://atomgit.com/openharmony/third_party_mindspore) |
| @@ -0,0 +1,251 @@ | |||
| 1 | +# GLiNER large-v2.5 MindSpore Lite 推理部署教程 | ||
| 2 | + | ||
| 3 | +本教程介绍如何将 [GLiNER large-v2.5](https://github.com/urchade/GLiNER) 导出为 ONNX 后转换为 MindSpore Lite MindIR,在 Atlas 300I Duo 上推理与测速。 | ||
| 4 | + | ||
| 5 | +GLiNER 是一种可指定任意标签的命名实体识别(NER)模型。本教程基于 GLiNER v0.2.27 的 `UniEncoderSpanGLiNER` 架构(DeBERTa-v3-large 主干 + 双向 LSTM + SpanMarkerV0 span 表示)。 | ||
| 6 | + | ||
| 7 | +--- | ||
| 8 | + | ||
| 9 | +## 1. 环境准备 | ||
| 10 | + | ||
| 11 | +### 依赖版本 | ||
| 12 | + | ||
| 13 | +| 软件包 | 版本 | | ||
| 14 | +| --- | --- | | ||
| 15 | +| Python | 3.11.15 | | ||
| 16 | +| torch | 2.10.0+cpu | | ||
| 17 | +| onnx | 1.19.1 | | ||
| 18 | +| numpy | 2.4.4 | | ||
| 19 | +| transformers | 4.57.0 | | ||
| 20 | +| gliner | 0.2.27 | | ||
| 21 | +| CANN | 8.5 | | ||
| 22 | +| mindspore-lite | 2.9.0 | | ||
| 23 | + | ||
| 24 | +```bash | ||
| 25 | +pip install torch==2.10.0 onnx==1.19.1 transformers==4.57.0 gliner==0.2.27 | ||
| 26 | +``` | ||
| 27 | + | ||
| 28 | +### 获取模型权重与源码 | ||
| 29 | + | ||
| 30 | +```bash | ||
| 31 | +# 模型源码 | ||
| 32 | +git clone https://github.com/urchade/GLiNER.git models/model_code/GLiNER | ||
| 33 | +pip install -e models/model_code/GLiNER | ||
| 34 | + | ||
| 35 | +# 模型权重(HuggingFace 下载) | ||
| 36 | +huggingface-cli download urchade/gliner_large-v2.5 \ | ||
| 37 | + --local-dir models/model_weight/gliner_large-v2.5 | ||
| 38 | +``` | ||
| 39 | + | ||
| 40 | +说明: | ||
| 41 | + | ||
| 42 | +- `MODEL_DIR`=`models/model_weight/gliner_large-v2.5`,包含 `pytorch_model.bin`、`gliner_config.json`、`tokenizer.json`、`spm.model` 等。 | ||
| 43 | +- 上游源码用于 import `gliner` 包;本目录脚本会在导出前对 `gliner` 内部函数做 monkey patch,把动态 shape 的 Python 控制流改造为可被 Ascend 接受的图。 | ||
| 44 | + | ||
| 45 | +--- | ||
| 46 | + | ||
| 47 | +## 2. 模型导出 ONNX | ||
| 48 | + | ||
| 49 | +### 导出命令 | ||
| 50 | + | ||
| 51 | +```bash | ||
| 52 | +cd mindspore-lite/examples/base_models/gliner_large-v2.5 | ||
| 53 | + | ||
| 54 | +python export_gliner_large_v2.5_onnx.py \ | ||
| 55 | + --model-dir models/model_weight/gliner_large-v2.5 \ | ||
| 56 | + --save-dir ./onnx \ | ||
| 57 | + --opset 17 | ||
| 58 | +``` | ||
| 59 | + | ||
| 60 | +### 参数说明 | ||
| 61 | + | ||
| 62 | +| 参数 | 说明 | 默认值 | | ||
| 63 | +| --- | --- | --- | | ||
| 64 | +| `--model-dir` | 权重目录(含 `pytorch_model.bin`、`gliner_config.json`、tokenizer 文件) | `models/model_weight/gliner_large-v2.5` | | ||
| 65 | +| `--save-dir` | ONNX 输出目录 | `./onnx` | | ||
| 66 | +| `--opset` | ONNX opset 版本 | `17` | | ||
| 67 | + | ||
| 68 | +### 产出文件 | ||
| 69 | + | ||
| 70 | +```text | ||
| 71 | +./onnx/ | ||
| 72 | +├── model.onnx # 单一 ONNX(含 DeBERTa + LSTM + SpanMarker) | ||
| 73 | +├── gliner_config.json # GLiNER 配置(max_width=12, ent/sep token 等) | ||
| 74 | +├── tokenizer.json # DeBERTa-v3 tokenizer | ||
| 75 | +├── spm.model # SentencePiece 模型 | ||
| 76 | +├── tokenizer_config.json | ||
| 77 | +├── special_tokens_map.json | ||
| 78 | +└── added_tokens.json | ||
| 79 | +``` | ||
| 80 | + | ||
| 81 | +### 导出注意事项(实际踩坑点) | ||
| 82 | + | ||
| 83 | +GLiNER 上游依赖三处 Ascend 不友好的实现,导出脚本在 `model.export_to_onnx(...)` 调用前对源码做了 monkey patch(均在 `export_gliner_large_v2.5_onnx.py` 中): | ||
| 84 | + | ||
| 85 | +1. **`_SmallOpLSTM` → 原生 `nn.LSTM`**:上游 `LstmSeq2SeqEncoder` 用 `_SmallOpLSTM`,其 `_run_direction` 是 Python `for t in range(seq_len)` 循环,JIT trace 会把 `seq_len` 烧成常量(dummy batch 是 5 words)。脚本从 checkpoint 中读取 `rnn.lstm.*_ih/_hh` 权重,重建为 `nn.LSTM` 并替换;同时 patch `LstmSeq2SeqEncoder.forward` 去掉 `lengths=` kwarg + 预分配静态 `h0/c0`。 | ||
| 86 | +2. **DeBERTa 的 `make_log_bucket_position`**:用 `@torch.jit.script` 包装,内部用 `torch.sign` 产生 Sign 算子,Ascend 不支持。脚本预计算 `(_REL_POS_MAX_SEQ × _REL_POS_MAX_SEQ)` 的 bucketed 相对位置矩阵作为常量,运行时只做切片。 | ||
| 87 | +3. **DeBERTa 的 `build_rpos` 与 `transpose_for_scores`**:`@torch.jit.script` 内的 Python `if` 会 trace 成 24 个 If 子图,Ascend 拒绝其中的 Range 子图;`transpose_for_scores` 用元组拼接,产生 rank 推断失败的 Concat。脚本把 `build_rpos` 替换为恒等函数,`transpose_for_scores` 用 Python int 常量重写。 | ||
| 88 | +4. **`extract_prompt_features`**:`.max()` 返回 ambiguous rank 的标量,与 span_idx 维度拼接时报 Concat rank 不匹配。脚本统一加 `.max().reshape(())` 强制 rank-0。 | ||
| 89 | + | ||
| 90 | +由于 dummy batch 默认用 `[person, organization, country]` 3 标签,导出的 ONNX 在 `num_classes=3` 维度上是动态的;若需要其他数量的标签,需要在 `gliner/model.py::_build_dummy_batch` 修改 `labels` 默认值,或直接在导出脚本中传入 `labels=` kwarg。 | ||
| 91 | + | ||
| 92 | +--- | ||
| 93 | + | ||
| 94 | +## 3. MindSpore Lite 转换(ONNX → MindIR) | ||
| 95 | + | ||
| 96 | +### 转换命令 | ||
| 97 | + | ||
| 98 | +说明:`converter_lite` 为 MindSpore Lite 版本包中提供的离线转换工具。 | ||
| 99 | + | ||
| 100 | +```bash | ||
| 101 | +converter_lite --fmk=ONNX \ | ||
| 102 | + --modelFile=./onnx/model.onnx \ | ||
| 103 | + --outputFile=./onnx/model \ | ||
| 104 | + --saveType=MINDIR \ | ||
| 105 | + --optimize=ascend_oriented \ | ||
| 106 | + --configFile=./gliner_large-v2.5.ini | ||
| 107 | +``` | ||
| 108 | + | ||
| 109 | +### 参数说明 | ||
| 110 | + | ||
| 111 | +| 参数 | 说明 | | ||
| 112 | +| --- | --- | | ||
| 113 | +| `--modelFile` | 输入 ONNX | | ||
| 114 | +| `--outputFile` | 输出前缀 | | ||
| 115 | +| `--optimize=ascend_oriented` | Ascend 定向优化 | | ||
| 116 | +| `--saveType=MINDIR` | 输出 MindIR | | ||
| 117 | +| `--configFile` | 配置文件(指定输入 dtype、固定 shape、precision mode 等) | | ||
| 118 | + | ||
| 119 | +### 配置文件 | ||
| 120 | + | ||
| 121 | +`gliner_large-v2.5.ini`(静态 shape,**生产推荐**): | ||
| 122 | + | ||
| 123 | +```ini | ||
| 124 | +[acl_build_options] | ||
| 125 | +input_format="ND" | ||
| 126 | +input_shape="input_ids:1,128;attention_mask:1,128;words_mask:1,128;text_lengths:1,1;span_idx:1,288,2;span_mask:1,288" | ||
| 127 | +``` | ||
| 128 | + | ||
| 129 | +固定 shape 说明(必须写清楚,否则推理侧无法对齐): | ||
| 130 | + | ||
| 131 | +| 输入 | 静态 shape | 含义 | | ||
| 132 | +| --- | --- | --- | | ||
| 133 | +| `input_ids` / `attention_mask` / `words_mask` | `(1, 128)` | 序列长度固定 128 | | ||
| 134 | +| `text_lengths` | `(1, 1)` | 实际 body words 数(运行时可变,1–24) | | ||
| 135 | +| `span_idx` | `(1, 288, 2)` | 24 body words × 12 max_width 的 span 网格 | | ||
| 136 | +| `span_mask` | `(1, 288)` | 仅 `s+k < text_lengths[0]` 的 span 为 True | | ||
| 137 | + | ||
| 138 | +**为什么必须静态 shape**:上游 `_SmallOpLSTM` 用 Python 循环展开,我们替换为原生 `nn.LSTM`,但原生 LSTM 在动态 seq_len 下输出 4D 张量 `[-1, 2, -1, 384]`(两个动态维度),Ascend 多 batch 编译器报 `Multi-batch not support middle dynamic shape`。尝试过预分配静态 `h0/c0` 与减少动态维度(`config_dyn_seq.ini`、`config.ini`)均失败,因此放弃动态 shape,转而使用静态 shape + padding/截断。 | ||
| 139 | + | ||
| 140 | +### 产出文件 | ||
| 141 | + | ||
| 142 | +```text | ||
| 143 | +./onnx/ | ||
| 144 | +└── model.mindir # ~1 GB,单文件(fp16 权重) | ||
| 145 | +``` | ||
| 146 | + | ||
| 147 | +执行日志: | ||
| 148 | + | ||
| 149 | +```log | ||
| 150 | +CONVERT RESULT SUCCESS:0 | ||
| 151 | +``` | ||
| 152 | + | ||
| 153 | +--- | ||
| 154 | + | ||
| 155 | +## 4. MindSpore Lite 推理 | ||
| 156 | + | ||
| 157 | +### 推理命令 | ||
| 158 | + | ||
| 159 | +```bash | ||
| 160 | +python infer_gliner_large_v2.5_mslite.py \ | ||
| 161 | + --model-dir models/model_weight/gliner_large-v2.5 \ | ||
| 162 | + --mindir-path ./onnx/model.mindir \ | ||
| 163 | + --text "Cristiano Ronaldo plays for Al-Nassr FC and captains Portugal." \ | ||
| 164 | + --labels person,organization,country \ | ||
| 165 | + --threshold 0.5 \ | ||
| 166 | + --device-id 0 | ||
| 167 | +``` | ||
| 168 | + | ||
| 169 | +### 参数说明 | ||
| 170 | + | ||
| 171 | +| 参数 | 说明 | 默认值 | | ||
| 172 | +| --- | --- | --- | | ||
| 173 | +| `--model-dir` | 权重目录(用于加载 tokenizer 与 `gliner_config.json`) | `models/model_weight/gliner_large-v2.5` | | ||
| 174 | +| `--mindir-path` | MindIR 文件 | `./onnx/model.mindir` | | ||
| 175 | +| `--text` | 输入文本 | 内置 3 条示例 | | ||
| 176 | +| `--text-file` | 每行一条文本的输入文件 | None | | ||
| 177 | +| `--labels` | 逗号分隔的实体标签(**必须为 3 个,与导出 dummy batch 一致**) | `person,organization,country` | | ||
| 178 | +| `--threshold` | sigmoid 置信度阈值 | `0.5` | | ||
| 179 | +| `--flat-ner` | 禁止 span 重叠(默认开启) | True | | ||
| 180 | +| `--device-id` | Ascend 设备 ID | `0` | | ||
| 181 | + | ||
| 182 | +### 执行日志 | ||
| 183 | + | ||
| 184 | +```log | ||
| 185 | +[infer] config: ent='<<ENT>>', sep='<<SEP>>', labels=['person', 'organization', 'country'] | ||
| 186 | +[infer] loading tokenizer from models/model_weight/gliner_large-v2.5 | ||
| 187 | +[infer] loading MindIR from ./onnx/model.mindir | ||
| 188 | +WARNING:root:Ascend custom operator path not found | ||
| 189 | + | ||
| 190 | +[infer] text: Cristiano Ronaldo Ronaldo dos Santos Aveiro plays for Al-Nassr FC and captains Portugal. | ||
| 191 | +[infer] words (13): ['Cristiano', 'Ronaldo', 'dos', 'Santos', 'Aveiro', 'plays', 'for', 'Al-Nassr', 'FC', 'and', 'captains', 'Portugal', '.'] | ||
| 192 | +[infer] seq_len: 128, logits shape: (1, 24, 12, 3) | ||
| 193 | + - 'Portugal' [country] score=0.9993 chars=(71, 79) | ||
| 194 | + - 'Al-Nassr FC' [organization] score=0.9944 chars=(46, 57) | ||
| 195 | + - 'Cristiano Ronaldo dos Santos Aveiro' [person] score=0.9910 chars=(0, 35) | ||
| 196 | +``` | ||
| 197 | + | ||
| 198 | +说明(ascend_oriented 固定 shape 约束): | ||
| 199 | + | ||
| 200 | +- 推理脚本固定 3 标签 `[person, organization, country]`,与导出 dummy batch 一致;传入其他数量的标签会报错。 | ||
| 201 | +- body words 数运行时可变(1–24),脚本自动截断长文本,并对 `input_ids/attention_mask/words_mask` 做 seq_len=128 padding。`text_lengths` 反映真实 body words 数,`span_mask` 仅置位前 `num_body_words × max_width` 个 span。 | ||
| 202 | +- 若需要不同的标签集,需要重新跑导出脚本(修改 `gliner/model.py::_build_dummy_batch` 的 `labels` 默认值,或在导出脚本中传 `labels=` kwarg),再重新转换 MindIR。 | ||
| 203 | + | ||
| 204 | +--- | ||
| 205 | + | ||
| 206 | +## 5. 性能数据 | ||
| 207 | + | ||
| 208 | +测试环境:Atlas 300I Duo | ||
| 209 | + | ||
| 210 | +固定文本 `"Cristiano Ronaldo dos Santos Aveiro plays for Al-Nassr FC and captains Portugal."`,3 标签,50 次平均(3 次 warmup 后): | ||
| 211 | + | ||
| 212 | +| 指标 | MindSpore Lite (Ascend fp16) | | ||
| 213 | +| --- | ---: | | ||
| 214 | +| 模型推理 | 15.83 ms | | ||
| 215 | +| 端到端(含预处理) | 16.49 ms | | ||
| 216 | +| **吞吐量** | **60.7 req/s** | | ||
| 217 | + | ||
| 218 | +--- | ||
| 219 | + | ||
| 220 | +## 6. 常见问题 | ||
| 221 | + | ||
| 222 | +1. 现象:`op[Expand], custom inputs shape [0] error!` | ||
| 223 | + - 原因:输入 `text_lengths` 为 0。 | ||
| 224 | + - 解决方案:保证至少有 1 个 body word;空文本会被跳过或加 placeholder。 | ||
| 225 | + | ||
| 226 | +2. 现象:converter 报 `Multi-batch not support middle dynamic shape. CurrentShape: [-1,-1,-1,-1]` | ||
| 227 | + - 原因:原生 `nn.LSTM` 在动态 seq_len 下输出 `[-1, 2, -1, 384]`,含 2 个动态维度,Ascend 多 batch 编译器拒绝。 | ||
| 228 | + - 解决方案:使用静态 shape(`gliner_large-v2.5.ini`),不要使用动态 shape 配置(`config.ini`/`config_dyn_seq.ini` 已废弃)。 | ||
| 229 | + | ||
| 230 | +3. 现象:`RuntimeError: Static MindIR has 3 classes baked in, but got N labels` | ||
| 231 | + - 原因:静态 MindIR 在转换时锁定了 `num_classes=3`(导出 dummy batch 用 3 标签)。 | ||
| 232 | + - 解决方案:传入 3 标签 `--labels person,organization,country`,或重新导出 ONNX 并重新转换。 | ||
| 233 | + | ||
| 234 | +4. 现象:converter 很慢且有大量 warning(`SetupParamInitSubGraph` / `tiling offset out of range`) | ||
| 235 | + - 原因:DeBERTa 主干层数深 + ascend_oriented 编译优化重。 | ||
| 236 | + - 解决方案:确认最终 `CONVERT RESULT SUCCESS:0`;转换约耗时 2–3 分钟,确保内存 ≥ 16 GB。 | ||
| 237 | + | ||
| 238 | +--- | ||
| 239 | + | ||
| 240 | +## 7. 参考资源 | ||
| 241 | + | ||
| 242 | +- 上游模型仓库:https://github.com/urchade/GLiNER | ||
| 243 | +- HuggingFace 权重:https://huggingface.co/urchade/gliner_large-v2.5 | ||
| 244 | +- MindSpore Lite 文档:https://www.mindspore.cn/lite | ||
| 245 | + | ||
| 246 | +--- | ||
| 247 | + | ||
| 248 | +## 8. 许可证 | ||
| 249 | + | ||
| 250 | +- 本目录脚本遵循 MindSpore Lite 仓库许可证要求。 | ||
| 251 | +- 上游 GLiNER 模型与代码以 Apache License 2.0 发布。 | ||
| @@ -0,0 +1,469 @@ | |||
| 1 | +#!/usr/bin/env python3 | ||
| 2 | +# Copyright 2026 Huawei Technologies Co., Ltd | ||
| 3 | +# | ||
| 4 | +# Licensed under the Apache License, Version 2.0 (the "License"); | ||
| 5 | +# you may not use this file except in compliance with the License. | ||
| 6 | +# You may obtain a copy of the License at | ||
| 7 | +# | ||
| 8 | +# http://www.apache.org/licenses/LICENSE-2.0 | ||
| 9 | +# | ||
| 10 | +# Unless required by applicable law or agreed to in writing, software | ||
| 11 | +# distributed under the License is distributed on an "AS IS" BASIS, | ||
| 12 | +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| 13 | +# See the License for the specific language governing permissions and | ||
| 14 | +# limitations under the License. | ||
| 15 | +# ============================================================================ | ||
| 16 | + | ||
| 17 | +"""Export gliner_large-v2.5 PyTorch model to ONNX. | ||
| 18 | + | ||
| 19 | +The upstream GLiNER package ships an ONNX-friendly ``_SmallOpLSTM`` (no | ||
| 20 | +``nn.LSTM`` / ``pack_padded_sequence``), but that implementation iterates the | ||
| 21 | +sequence with a Python ``for t in range(seq_len)`` loop. JIT tracing unrolls | ||
| 22 | +that loop to the dummy batch's word count, which then becomes a hard cap on | ||
| 23 | +the runtime sequence length: any input longer than the dummy's word count is | ||
| 24 | +silently truncated, and shorter inputs leak stale state across the unused | ||
| 25 | +positions. To get a model that actually honors the dynamic ``sequence_length`` | ||
| 26 | +axis, we swap ``_SmallOpLSTM`` for a native ``nn.LSTM`` (C++ implementation, | ||
| 27 | +no Python loop, traces cleanly to a single dynamic-shape ONNX subgraph). | ||
| 28 | + | ||
| 29 | +The published ``gliner-community/gliner_large-v2.5`` checkpoint was saved | ||
| 30 | +with native ``nn.LSTM`` weights (``weight_ih_l0`` / ``weight_hh_l0`` etc.), | ||
| 31 | +so swapping in ``nn.LSTM`` also removes the need for any weight-name | ||
| 32 | +remapping — the checkpoint keys line up directly with the new module. | ||
| 33 | + | ||
| 34 | +On top of that fix, three export blockers remain: | ||
| 35 | + | ||
| 36 | +1. ``UniEncoderSpanModel.forward`` uses ``torch.einsum("BLKD,BCD->BLKC", ...)`` | ||
| 37 | + for span-vs-prompt scoring. Einsum is a common conversion blocker on Ascend, | ||
| 38 | + so we rewrite it as ``matmul`` + ``reshape`` ahead of time. | ||
| 39 | +2. ``_fit_length`` uses Python ``if target_len == L:`` / ``if target_len > L:`` | ||
| 40 | + to pick between slicing and padding. JIT tracing bakes the branch taken at | ||
| 41 | + trace time, so when the dummy batch has fewer words than a real batch the | ||
| 42 | + graph silently picks the wrong branch and the downstream gather fails with | ||
| 43 | + out-of-range indices. We replace it with a branchless pre-pad + dynamic | ||
| 44 | + slice that always works regardless of runtime shape. | ||
| 45 | +3. DeBERTa-v2's disentangled attention uses a ``@torch.jit.script``-decorated | ||
| 46 | + ``make_log_bucket_position``. When traced, that scripted function emits one | ||
| 47 | + ONNX ``If`` subgraph per attention layer (24 in DeBERTa-large) plus a | ||
| 48 | + ``Sign`` op — both rejected by MSLite's Ascend converter. We precompute the | ||
| 49 | + bucketed relative-position matrix once at export time as a plain tensor and | ||
| 50 | + slice it at runtime, eliminating every ``If``/``Sign`` from the graph. | ||
| 51 | + | ||
| 52 | +The script then delegates to ``GLiNER.export_to_onnx`` which already wires up | ||
| 53 | +the dummy batch, the dynamic axes, and the I/O spec. | ||
| 54 | +""" | ||
| 55 | + | ||
| 56 | +import argparse | ||
| 57 | +import os | ||
| 58 | +import sys | ||
| 59 | +from pathlib import Path | ||
| 60 | + | ||
| 61 | +import torch | ||
| 62 | +from torch import nn | ||
| 63 | +from gliner import GLiNER | ||
| 64 | +from gliner.modeling.base import GLiNERBaseOutput, UniEncoderSpanModel | ||
| 65 | +from gliner.modeling.layers import LstmSeq2SeqEncoder | ||
| 66 | +import gliner.modeling.base as gbase | ||
| 67 | +import gliner.modeling.utils as gutils | ||
| 68 | +import transformers.models.deberta_v2.modeling_deberta_v2 as dv2 | ||
| 69 | +from transformers.models.deberta_v2.modeling_deberta_v2 import make_log_bucket_position | ||
| 70 | + | ||
| 71 | +DEFAULT_MODEL_DIR = "gliner_large-v2.5" | ||
| 72 | +DEFAULT_SAVE_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "onnx") | ||
| 73 | + | ||
| 74 | +# Idempotency registry: each _patch_* helper checks ``key in _PATCHED`` so | ||
| 75 | +# re-invoking export in the same process is a no-op instead of re-wrapping. | ||
| 76 | +_PATCHED = set() | ||
| 77 | + | ||
| 78 | +# Safe upper bound for words_embedding length after _fit_length. | ||
| 79 | +# max_len=768 in gliner_config.json, so target_W <= ~768 + max_width. | ||
| 80 | +_FIT_LENGTH_MAX_PAD = 1024 | ||
| 81 | + | ||
| 82 | +# Safe upper bound for the relative-position buffer. The gliner_config has | ||
| 83 | +# max_len=768 and the prompt adds ~10 tokens, so 1024 is comfortably larger. | ||
| 84 | +_REL_POS_MAX_SEQ = 1024 | ||
| 85 | +# Single-element mutable holder so _patch_relative_position can populate the | ||
| 86 | +# buffer without a ``global`` declaration. | ||
| 87 | +_REL_POS_BUFFER_HOLDER = [None] | ||
| 88 | + | ||
| 89 | + | ||
| 90 | +def _fit_length_dynamic(embedding, mask, target_len): | ||
| 91 | + """Branchless, ONNX-friendly replacement for ``UniEncoderSpanModel._fit_length``. | ||
| 92 | + | ||
| 93 | + The stock implementation picks between slicing and padding with a Python | ||
| 94 | + ``if``, which JIT tracing bakes as a single branch. When the runtime | ||
| 95 | + ``target_len`` falls in the other branch the graph silently produces the | ||
| 96 | + wrong shape and the downstream ``GatherElements`` fails. | ||
| 97 | + | ||
| 98 | + We always pre-pad with ``_FIT_LENGTH_MAX_PAD`` zeros along dim=1 and then | ||
| 99 | + slice to ``target_len``. ``target_len`` is computed at runtime from | ||
| 100 | + ``span_idx.size(1) // max_width`` and propagates as a dynamic value | ||
| 101 | + through the ONNX graph, so the slice adjusts correctly for any input. | ||
| 102 | + """ | ||
| 103 | + b = embedding.size(0) | ||
| 104 | + d = embedding.size(-1) | ||
| 105 | + zeros_emb = torch.zeros(b, _FIT_LENGTH_MAX_PAD, d, dtype=embedding.dtype, device=embedding.device) | ||
| 106 | + emb_padded = torch.cat([embedding, zeros_emb], dim=1) | ||
| 107 | + zeros_mask = torch.zeros(b, _FIT_LENGTH_MAX_PAD, dtype=mask.dtype, device=mask.device) | ||
| 108 | + mask_padded = torch.cat([mask, zeros_mask], dim=1) | ||
| 109 | + return emb_padded[:, :target_len], mask_padded[:, :target_len] | ||
| 110 | + | ||
| 111 | + | ||
| 112 | +def _patch_lstm_seq2seq_forward() -> int: | ||
| 113 | + """Monkey-patch ``LstmSeq2SeqEncoder.forward`` to call native ``nn.LSTM``. | ||
| 114 | + | ||
| 115 | + The stock implementation passes ``lengths=lengths`` to ``self.lstm`` (which | ||
| 116 | + is a ``_SmallOpLSTM`` accepting that kwarg) and then slices | ||
| 117 | + ``output[:, :max_len]`` with ``max_len = int(lengths.max().item())``. After | ||
| 118 | + ``_swap_in_native_lstm`` swaps in a native ``nn.LSTM`` (which has no | ||
| 119 | + ``lengths`` parameter), we must drop the kwarg; and the | ||
| 120 | + ``int(...max().item())`` slice must be replaced with ``x.size(1)`` so JIT | ||
| 121 | + tracing keeps the sequence dimension symbolic instead of baking the dummy | ||
| 122 | + batch's max word count as a constant (which would silently truncate | ||
| 123 | + longer real inputs at inference time). | ||
| 124 | + | ||
| 125 | + We also pre-allocate ``h0`` / ``c0`` as static zero tensors sized to the | ||
| 126 | + native ``nn.LSTM``'s expected ``(num_layers * num_directions, batch=1, | ||
| 127 | + hidden_size)`` shape. With ``hidden=None`` PyTorch traces the zero init as | ||
| 128 | + a ``ConstantOfShape`` whose shape is itself derived from the LSTM input | ||
| 129 | + via ``Shape``/``Gather``/``Concat`` — and Ascend's multi-batch compiler | ||
| 130 | + rejects the resulting 4-dynamic-dim LSTM output. Pre-allocated constant | ||
| 131 | + initial states eliminate that shape-construction chain entirely. | ||
| 132 | + """ | ||
| 133 | + if "lstm_forward" in _PATCHED: | ||
| 134 | + return 0 | ||
| 135 | + | ||
| 136 | + def patched_forward(self, x, mask, hidden=None, lengths=None): | ||
| 137 | + del mask, lengths # unused; native nn.LSTM doesn't accept lengths | ||
| 138 | + if hidden is None: | ||
| 139 | + lstm = self.lstm | ||
| 140 | + n_dirs = 2 if lstm.bidirectional else 1 | ||
| 141 | + n_layers = lstm.num_layers | ||
| 142 | + h = torch.zeros(n_layers * n_dirs, x.size(0), lstm.hidden_size, | ||
| 143 | + dtype=x.dtype, device=x.device) | ||
| 144 | + c = torch.zeros(n_layers * n_dirs, x.size(0), lstm.hidden_size, | ||
| 145 | + dtype=x.dtype, device=x.device) | ||
| 146 | + hidden = (h, c) | ||
| 147 | + output, _ = self.lstm(x, hidden) | ||
| 148 | + return output[:, : x.size(1)] | ||
| 149 | + | ||
| 150 | + LstmSeq2SeqEncoder.forward = patched_forward | ||
| 151 | + _PATCHED.add("lstm_forward") | ||
| 152 | + return 1 | ||
| 153 | + | ||
| 154 | + | ||
| 155 | +def _precompute_rel_pos_buffer(bucket_size: int, max_position: int) -> torch.Tensor: | ||
| 156 | + """Precompute ``make_log_bucket_position`` output for all rel_pos in [-N+1, N-1]. | ||
| 157 | + | ||
| 158 | + DeBERTa-v2's disentangled attention uses a JIT-scripted | ||
| 159 | + ``make_log_bucket_position`` that, when traced, emits one ONNX ``If`` | ||
| 160 | + subgraph per attention layer (24 in DeBERTa-large) plus a ``Sign`` op. | ||
| 161 | + MSLite's Ascend converter rejects both. We sidestep this entirely by | ||
| 162 | + computing the bucketed relative-position matrix ONCE at export time as a | ||
| 163 | + plain tensor — the bucketing depends only on ``q - k`` and the layer's | ||
| 164 | + bucket_size / max_position, not on token content — and then slicing the | ||
| 165 | + cached buffer at runtime to the actual ``[query_size, key_size]`` shape. | ||
| 166 | + """ | ||
| 167 | + n = _REL_POS_MAX_SEQ | ||
| 168 | + q_ids = torch.arange(n, dtype=torch.long) | ||
| 169 | + k_ids = torch.arange(n, dtype=torch.long) | ||
| 170 | + rel_pos = q_ids[:, None] - k_ids[None, :] | ||
| 171 | + return make_log_bucket_position(rel_pos, bucket_size, max_position).to(torch.long) | ||
| 172 | + | ||
| 173 | + | ||
| 174 | +def _find_disentangled_attention(model): | ||
| 175 | + """Return the first ``DisentangledSelfAttention`` module, or None.""" | ||
| 176 | + for module in model.modules(): | ||
| 177 | + if module.__class__.__name__ == "DisentangledSelfAttention": | ||
| 178 | + return module | ||
| 179 | + return None | ||
| 180 | + | ||
| 181 | + | ||
| 182 | +def _patch_relative_position(attn_module) -> int: | ||
| 183 | + """Replace ``build_relative_position`` with a slice into a precomputed buffer.""" | ||
| 184 | + if "rel_pos" in _PATCHED: | ||
| 185 | + return 0 | ||
| 186 | + if attn_module is None: | ||
| 187 | + return 0 | ||
| 188 | + | ||
| 189 | + _REL_POS_BUFFER_HOLDER[0] = _precompute_rel_pos_buffer( | ||
| 190 | + attn_module.position_buckets, attn_module.max_relative_positions | ||
| 191 | + ) | ||
| 192 | + | ||
| 193 | + def patched_build_relative_position(query_layer, key_layer, bucket_size=-1, max_position=-1): | ||
| 194 | + del bucket_size, max_position # already baked into the buffer | ||
| 195 | + query_size = query_layer.size(-2) | ||
| 196 | + key_size = key_layer.size(-2) | ||
| 197 | + return _REL_POS_BUFFER_HOLDER[0][:query_size, :key_size].unsqueeze(0).to(query_layer.device) | ||
| 198 | + | ||
| 199 | + dv2.build_relative_position = patched_build_relative_position | ||
| 200 | + _PATCHED.add("rel_pos") | ||
| 201 | + return 1 | ||
| 202 | + | ||
| 203 | + | ||
| 204 | +def _patch_build_rpos() -> int: | ||
| 205 | + """Replace JIT-scripted ``build_rpos`` with an identity (key_size == query_size in self-attn).""" | ||
| 206 | + if "build_rpos" in _PATCHED: | ||
| 207 | + return 0 | ||
| 208 | + | ||
| 209 | + def patched_build_rpos(query_layer, key_layer, relative_pos, position_buckets=-1, max_relative_positions=-1): | ||
| 210 | + del query_layer, key_layer, position_buckets, max_relative_positions | ||
| 211 | + return relative_pos | ||
| 212 | + | ||
| 213 | + dv2.build_rpos = patched_build_rpos | ||
| 214 | + _PATCHED.add("build_rpos") | ||
| 215 | + return 1 | ||
| 216 | + | ||
| 217 | + | ||
| 218 | +def _patch_transpose_for_scores(attn_module) -> int: | ||
| 219 | + """Replace ``transpose_for_scores`` with a Python-int-only reshape.""" | ||
| 220 | + if "transpose_for_scores" in _PATCHED: | ||
| 221 | + return 0 | ||
| 222 | + if attn_module is None: | ||
| 223 | + return 0 | ||
| 224 | + heads = attn_module.num_attention_heads | ||
| 225 | + hidden_size = getattr(attn_module, "all_head_size", None) | ||
| 226 | + if hidden_size is None: | ||
| 227 | + hidden_size = getattr(getattr(attn_module, "query_proj", None), "in_features", None) | ||
| 228 | + if hidden_size is None: | ||
| 229 | + return 0 | ||
| 230 | + head_dim_py = hidden_size // heads | ||
| 231 | + | ||
| 232 | + def patched_transpose_for_scores(self, x, attention_heads): | ||
| 233 | + del self, attention_heads # captured at patch time; matches heads_py | ||
| 234 | + # (B, S, hidden) → (B, S, heads, head_dim); -1 absorbs S. | ||
| 235 | + x = x.view(-1, x.size(1), heads, head_dim_py) | ||
| 236 | + x = x.permute(0, 2, 1, 3).contiguous() | ||
| 237 | + # (B, heads, S, head_dim) → (-1, S, head_dim) flattens B*heads. | ||
| 238 | + return x.view(-1, x.size(2), head_dim_py) | ||
| 239 | + | ||
| 240 | + dv2.DisentangledSelfAttention.transpose_for_scores = patched_transpose_for_scores | ||
| 241 | + _PATCHED.add("transpose_for_scores") | ||
| 242 | + return 1 | ||
| 243 | + | ||
| 244 | + | ||
| 245 | +def _patched_extract_prompt_features(class_token_index, token_embeds, input_ids, attention_mask, | ||
| 246 | + batch_size, embed_dim, embed_ent_token=True): | ||
| 247 | + """Force ``.max()`` scalars to rank-0 via ``.reshape(())`` so downstream Concat infers.""" | ||
| 248 | + class_token_mask = input_ids == class_token_index | ||
| 249 | + num_class_tokens = torch.sum(class_token_mask, dim=-1, keepdim=True) | ||
| 250 | + max_embed_dim = num_class_tokens.max().reshape(()) | ||
| 251 | + aranged_class_idx = torch.arange(max_embed_dim, dtype=attention_mask.dtype, device=token_embeds.device).expand( | ||
| 252 | + batch_size, -1 | ||
| 253 | + ) | ||
| 254 | + batch_indices, target_class_idx = torch.where(aranged_class_idx < num_class_tokens) | ||
| 255 | + _, class_indices = torch.where(class_token_mask) | ||
| 256 | + if not embed_ent_token: | ||
| 257 | + class_indices = class_indices + 1 | ||
| 258 | + prompts_embedding = torch.zeros( | ||
| 259 | + batch_size, max_embed_dim, embed_dim, dtype=token_embeds.dtype, device=token_embeds.device | ||
| 260 | + ) | ||
| 261 | + prompts_embedding[batch_indices, target_class_idx] = token_embeds[batch_indices, class_indices] | ||
| 262 | + prompts_embedding_mask = (aranged_class_idx < num_class_tokens).to(attention_mask.dtype) | ||
| 263 | + return prompts_embedding, prompts_embedding_mask | ||
| 264 | + | ||
| 265 | + | ||
| 266 | +def _patched_extract_prompt_features_and_word_embeddings(class_token_index, token_embeds, input_ids, | ||
| 267 | + attention_mask, text_lengths, words_mask, | ||
| 268 | + embed_ent_token=True, **kwargs): | ||
| 269 | + """Same rank-0 fix applied to ``extract_prompt_features_and_word_embeddings``.""" | ||
| 270 | + del kwargs | ||
| 271 | + batch_size, _, embed_dim = token_embeds.shape | ||
| 272 | + max_text_length = text_lengths.max().reshape(()) | ||
| 273 | + prompts_embedding, prompts_embedding_mask = _patched_extract_prompt_features( | ||
| 274 | + class_token_index, token_embeds, input_ids, attention_mask, batch_size, embed_dim, embed_ent_token | ||
| 275 | + ) | ||
| 276 | + words_embedding, mask = gutils.extract_word_embeddings( | ||
| 277 | + token_embeds, words_mask, attention_mask, batch_size, max_text_length, embed_dim, text_lengths | ||
| 278 | + ) | ||
| 279 | + return prompts_embedding, prompts_embedding_mask, words_embedding, mask | ||
| 280 | + | ||
| 281 | + | ||
| 282 | +def _patch_extract_prompt_features() -> int: | ||
| 283 | + """Patch prompt-feature extractors in both ``gliner.modeling.utils`` and ``base``.""" | ||
| 284 | + if "extract_prompt_features" in _PATCHED: | ||
| 285 | + return 0 | ||
| 286 | + | ||
| 287 | + gutils.extract_prompt_features = _patched_extract_prompt_features | ||
| 288 | + gutils.extract_prompt_features_and_word_embeddings = _patched_extract_prompt_features_and_word_embeddings | ||
| 289 | + gbase.extract_prompt_features = _patched_extract_prompt_features | ||
| 290 | + gbase.extract_prompt_features_and_word_embeddings = _patched_extract_prompt_features_and_word_embeddings | ||
| 291 | + _PATCHED.add("extract_prompt_features") | ||
| 292 | + return 1 | ||
| 293 | + | ||
| 294 | + | ||
| 295 | +def _patch_deberta_build_relative_position(model) -> int: | ||
| 296 | + """Apply all four DeBERTa-side patches needed for Ascend conversion. | ||
| 297 | + | ||
| 298 | + Patches: | ||
| 299 | + 1. ``build_relative_position`` → slice into precomputed bucket buffer (Sign op) | ||
| 300 | + 2. ``build_rpos`` → identity (If/Range ops in JIT-scripted branch) | ||
| 301 | + 3. ``transpose_for_scores`` → Python-int-only reshape (Concat rank mismatch) | ||
| 302 | + 4. ``extract_prompt_features`` → ``.max().reshape(())`` for rank-0 scalars | ||
| 303 | + """ | ||
| 304 | + attn_module = _find_disentangled_attention(model) | ||
| 305 | + n1 = _patch_relative_position(attn_module) | ||
| 306 | + n2 = _patch_build_rpos() | ||
| 307 | + n3 = _patch_transpose_for_scores(attn_module) | ||
| 308 | + n4 = _patch_extract_prompt_features() | ||
| 309 | + return n1 + n2 + n3 + n4 | ||
| 310 | + | ||
| 311 | + | ||
| 312 | +def _patch_uni_encoder_forward() -> int: | ||
| 313 | + """Monkey-patch ``UniEncoderSpanModel.forward`` for clean ONNX export. | ||
| 314 | + | ||
| 315 | + Two changes vs. upstream: | ||
| 316 | + - Replace ``einsum("BLKD,BCD->BLKC", ...)`` with ``matmul`` after a | ||
| 317 | + reshape/transpose. Einsum is a common Ascend conversion blocker. | ||
| 318 | + - Replace ``_fit_length`` with ``_fit_length_dynamic`` so the model | ||
| 319 | + handles any runtime ``span_idx`` shape (see its docstring for why). | ||
| 320 | + """ | ||
| 321 | + if "uni_encoder_forward" in _PATCHED: | ||
| 322 | + return 0 | ||
| 323 | + | ||
| 324 | + def patched_forward(self, input_ids=None, attention_mask=None, words_embedding=None, | ||
| 325 | + mask=None, prompts_embedding=None, prompts_embedding_mask=None, | ||
| 326 | + words_mask=None, text_lengths=None, span_idx=None, span_mask=None, | ||
| 327 | + labels=None, **kwargs): | ||
| 328 | + del words_embedding, mask, prompts_embedding, prompts_embedding_mask, labels, kwargs | ||
| 329 | + | ||
| 330 | + prompts_embedding, prompts_embedding_mask, words_embedding, mask = self.get_representations( | ||
| 331 | + input_ids, attention_mask, text_lengths, words_mask | ||
| 332 | + ) | ||
| 333 | + | ||
| 334 | + target_w = span_idx.size(1) // self.config.max_width | ||
| 335 | + words_embedding, mask = _fit_length_dynamic(words_embedding, mask, target_w) | ||
| 336 | + | ||
| 337 | + span_idx = span_idx * span_mask.unsqueeze(-1) | ||
| 338 | + span_rep = self.span_rep_layer(words_embedding, span_idx) | ||
| 339 | + | ||
| 340 | + # During inference labels is None, so target_c == prompts_embedding.size(1) | ||
| 341 | + # and _fit_length is a no-op. Skip the call entirely to avoid baking a | ||
| 342 | + # Python int into the graph. | ||
| 343 | + prompts_embedding = self.prompt_rep_layer(prompts_embedding) | ||
| 344 | + | ||
| 345 | + b, length, k, d = span_rep.shape | ||
| 346 | + c = prompts_embedding.size(1) | ||
| 347 | + span_rep_flat = span_rep.reshape(b, length * k, d) | ||
| 348 | + prompts_t = prompts_embedding.transpose(1, 2) | ||
| 349 | + scores = torch.matmul(span_rep_flat, prompts_t).reshape(b, length, k, c) | ||
| 350 | + | ||
| 351 | + return GLiNERBaseOutput( | ||
| 352 | + logits=scores, | ||
| 353 | + loss=None, | ||
| 354 | + prompts_embedding=prompts_embedding, | ||
| 355 | + prompts_embedding_mask=prompts_embedding_mask, | ||
| 356 | + words_embedding=words_embedding, | ||
| 357 | + mask=mask, | ||
| 358 | + ) | ||
| 359 | + | ||
| 360 | + UniEncoderSpanModel.forward = patched_forward | ||
| 361 | + _PATCHED.add("uni_encoder_forward") | ||
| 362 | + return 1 | ||
| 363 | + | ||
| 364 | + | ||
| 365 | +def _swap_in_native_lstm(model, checkpoint_dir) -> int: | ||
| 366 | + """Replace ``_SmallOpLSTM`` with a native ``nn.LSTM`` loaded from the checkpoint. | ||
| 367 | + | ||
| 368 | + The upstream ``LstmSeq2SeqEncoder`` ships with ``_SmallOpLSTM`` — a Python | ||
| 369 | + loop implementation of the BiLSTM. JIT tracing unrolls that loop to the | ||
| 370 | + dummy batch's word count, which then becomes a hard cap on the runtime | ||
| 371 | + sequence length (longer inputs are silently truncated at inference time). | ||
| 372 | + A native ``nn.LSTM`` has no Python loop and traces cleanly to a single | ||
| 373 | + dynamic-shape ONNX subgraph. | ||
| 374 | + | ||
| 375 | + The published ``gliner_large-v2.5`` checkpoint was saved with native | ||
| 376 | + ``nn.LSTM`` weights (``weight_ih_l0`` / ``weight_hh_l0`` etc.), so we can | ||
| 377 | + construct an ``nn.LSTM`` with the same hyperparameters as the original | ||
| 378 | + ``_SmallOpLSTM``, load the checkpoint keys directly (no name remapping), | ||
| 379 | + and substitute it in place via ``rnn.lstm = nn_lstm``. | ||
| 380 | + | ||
| 381 | + Numerically this is equivalent to ``_SmallOpLSTM`` at runtime because the | ||
| 382 | + collator pads every batch to its own ``lengths.max()``, so | ||
| 383 | + ``x.size(1) == lengths.max()`` and there is no real padding in the input — | ||
| 384 | + every position is active, which is exactly the case where the masked | ||
| 385 | + ``_SmallOpLSTM`` and the unmasked native ``nn.LSTM`` agree. | ||
| 386 | + """ | ||
| 387 | + rnn = getattr(model.model, "rnn", None) | ||
| 388 | + if rnn is None or not hasattr(rnn, "lstm"): | ||
| 389 | + return 0 | ||
| 390 | + small_lstm = rnn.lstm | ||
| 391 | + if isinstance(small_lstm, nn.LSTM): | ||
| 392 | + return 0 # already swapped (e.g. by a previous call in the same process) | ||
| 393 | + | ||
| 394 | + state_path = _find_checkpoint(checkpoint_dir) | ||
| 395 | + if state_path is None: | ||
| 396 | + return 0 | ||
| 397 | + state_dict = torch.load(state_path, map_location="cpu", weights_only=True) | ||
| 398 | + nn_keys = [k for k in state_dict if k.startswith("rnn.lstm.") and "_l" in k] | ||
| 399 | + if not nn_keys: | ||
| 400 | + return 0 | ||
| 401 | + | ||
| 402 | + nn_lstm = nn.LSTM( | ||
| 403 | + input_size=small_lstm.input_size, | ||
| 404 | + hidden_size=small_lstm.hidden_size, | ||
| 405 | + num_layers=small_lstm.num_layers, | ||
| 406 | + bidirectional=small_lstm.bidirectional, | ||
| 407 | + batch_first=True, | ||
| 408 | + ) | ||
| 409 | + nn_state = {k[len("rnn.lstm."):]: state_dict[k] for k in nn_keys} | ||
| 410 | + nn_lstm.load_state_dict(nn_state, strict=True) | ||
| 411 | + nn_lstm.eval() | ||
| 412 | + rnn.lstm = nn_lstm # in-place swap; _SmallOpLSTM no longer reachable | ||
| 413 | + return len(nn_state) | ||
| 414 | + | ||
| 415 | + | ||
| 416 | +def _find_checkpoint(checkpoint_dir) -> Path: | ||
| 417 | + """Return the .bin/.safetensors checkpoint path inside ``checkpoint_dir``.""" | ||
| 418 | + for name in ("pytorch_model.bin", "model.safetensors"): | ||
| 419 | + path = Path(checkpoint_dir) / name | ||
| 420 | + if path.exists(): | ||
| 421 | + return path | ||
| 422 | + return None | ||
| 423 | + | ||
| 424 | + | ||
| 425 | +def parse_args() -> argparse.Namespace: | ||
| 426 | + """Parse command-line arguments.""" | ||
| 427 | + parser = argparse.ArgumentParser(description="Export gliner_large-v2.5 to ONNX") | ||
| 428 | + parser.add_argument("--model-dir", default=DEFAULT_MODEL_DIR, help="Path to gliner checkpoint") | ||
| 429 | + parser.add_argument("--save-dir", default=DEFAULT_SAVE_DIR, help="Output directory for ONNX") | ||
| 430 | + parser.add_argument("--opset", type=int, default=17, help="ONNX opset version") | ||
| 431 | + return parser.parse_args() | ||
| 432 | + | ||
| 433 | + | ||
| 434 | +def main() -> None: | ||
| 435 | + """Load GLiNER, swap in native nn.LSTM, patch einsum/_fit_length, export to ONNX.""" | ||
| 436 | + args = parse_args() | ||
| 437 | + save_dir = Path(args.save_dir) | ||
| 438 | + save_dir.mkdir(parents=True, exist_ok=True) | ||
| 439 | + | ||
| 440 | + print(f"[export] loading GLiNER from {args.model_dir}", flush=True) | ||
| 441 | + model = GLiNER.from_pretrained(args.model_dir, load_tokenizer=True) | ||
| 442 | + model.eval() | ||
| 443 | + | ||
| 444 | + n_lstm_weights = _swap_in_native_lstm(model, args.model_dir) | ||
| 445 | + print(f"[export] swapped in native nn.LSTM (loaded {n_lstm_weights} tensors)", flush=True) | ||
| 446 | + | ||
| 447 | + n_patched = _patch_uni_encoder_forward() | ||
| 448 | + print(f"[export] patched UniEncoderSpanModel.forward (matmul + dynamic fit_length): {n_patched}", | ||
| 449 | + flush=True) | ||
| 450 | + | ||
| 451 | + n_lstm_patched = _patch_lstm_seq2seq_forward() | ||
| 452 | + print(f"[export] patched LstmSeq2SeqEncoder.forward (native nn.LSTM call, symbolic max_len): " | ||
| 453 | + f"{n_lstm_patched}", flush=True) | ||
| 454 | + | ||
| 455 | + n_rel_pos_patched = _patch_deberta_build_relative_position(model) | ||
| 456 | + print(f"[export] patched DeBERTa build_relative_position (precomputed bucket buffer): " | ||
| 457 | + f"{n_rel_pos_patched}", flush=True) | ||
| 458 | + | ||
| 459 | + print(f"[export] exporting ONNX to {save_dir}/model.onnx (opset={args.opset})", flush=True) | ||
| 460 | + model.export_to_onnx( | ||
| 461 | + save_dir=save_dir, | ||
| 462 | + onnx_filename="model.onnx", | ||
| 463 | + opset=args.opset, | ||
| 464 | + ) | ||
| 465 | + print("[export] done. files:", sorted(os.listdir(save_dir)), flush=True) | ||
| 466 | + | ||
| 467 | + | ||
| 468 | +if __name__ == "__main__": | ||
| 469 | + sys.exit(main()) | ||
| @@ -0,0 +1,3 @@ | |||
| 1 | +[acl_build_options] | ||
| 2 | +input_format="ND" | ||
| 3 | +input_shape="input_ids:1,128;attention_mask:1,128;words_mask:1,128;text_lengths:1,1;span_idx:1,288,2;span_mask:1,288" | ||
| @@ -0,0 +1,273 @@ | |||
| 1 | +#!/usr/bin/env python3 | ||
| 2 | +# Copyright 2026 Huawei Technologies Co., Ltd | ||
| 3 | +# | ||
| 4 | +# Licensed under the Apache License, Version 2.0 (the "License"); | ||
| 5 | +# you may not use this file except in compliance with the License. | ||
| 6 | +# You may obtain a copy of the License at | ||
| 7 | +# | ||
| 8 | +# http://www.apache.org/licenses/LICENSE-2.0 | ||
| 9 | +# | ||
| 10 | +# Unless required by applicable law or agreed to in writing, software | ||
| 11 | +# distributed under the License is distributed on an "AS IS" BASIS, | ||
| 12 | +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| 13 | +# See the License for the specific language governing permissions and | ||
| 14 | +# limitations under the License. | ||
| 15 | +# ============================================================================ | ||
| 16 | + | ||
| 17 | +"""MindSpore Lite inference script for gliner_large-v2.5. | ||
| 18 | + | ||
| 19 | +Same end-to-end NER pipeline for MindSpore Lite running on Ascend. Static-shape MindIR is used | ||
| 20 | +because dynamic seq_len + native nn.LSTM is rejected by the Ascend | ||
| 21 | +multi-batch compiler. The static shape is: | ||
| 22 | + | ||
| 23 | + input_ids/attention_mask/words_mask : (1, 128) | ||
| 24 | + text_lengths : (1, 1) | ||
| 25 | + span_idx : (1, 288, 2) # 24 words * 12 widths | ||
| 26 | + span_mask : (1, 288) | ||
| 27 | + logits : (1, 24, 12, 3) | ||
| 28 | + | ||
| 29 | +The label set is baked into the MindIR (the ONNX export dummy batch used | ||
| 30 | +``[person, organization, country]``). To use a different label set, re-export | ||
| 31 | +the ONNX with ``--labels`` (in the upstream GLiNER exporter) and re-convert. | ||
| 32 | + | ||
| 33 | +No ``torch`` import: pure numpy + transformers + mindspore_lite. | ||
| 34 | +""" | ||
| 35 | + | ||
| 36 | +import argparse | ||
| 37 | +import json | ||
| 38 | +import re | ||
| 39 | +from pathlib import Path | ||
| 40 | + | ||
| 41 | +import numpy as np | ||
| 42 | +import mindspore_lite as mslite | ||
| 43 | +from transformers import AutoTokenizer | ||
| 44 | + | ||
| 45 | +# Fixed shapes baked into the MindIR. | ||
| 46 | +SEQ_LEN = 128 | ||
| 47 | +NUM_WORDS = 24 | ||
| 48 | +MAX_WIDTH = 12 | ||
| 49 | +NUM_CLASSES = 3 | ||
| 50 | + | ||
| 51 | +# Label set must match the labels used at ONNX export time. | ||
| 52 | +DEFAULT_LABELS = ["person", "organization", "country"] | ||
| 53 | + | ||
| 54 | +_WHITESPACE_PATTERN = re.compile(r"\w+(?:[-_]\w+)*|\S") | ||
| 55 | + | ||
| 56 | + | ||
| 57 | +def _split_words(text): | ||
| 58 | + """Split text into (token, char_start, char_end) using GLiNER's whitespace rule.""" | ||
| 59 | + return [(m.group(), m.start(), m.end()) for m in _WHITESPACE_PATTERN.finditer(text)] | ||
| 60 | + | ||
| 61 | + | ||
| 62 | +def _build_prompt(labels, ent_token, sep_token): | ||
| 63 | + """Build the GLiNER prompt: [ENT, label1, ENT, label2, ..., SEP].""" | ||
| 64 | + prompt = [] | ||
| 65 | + for label in labels: | ||
| 66 | + prompt.append(ent_token) | ||
| 67 | + prompt.append(label) | ||
| 68 | + prompt.append(sep_token) | ||
| 69 | + return prompt | ||
| 70 | + | ||
| 71 | + | ||
| 72 | +def _tokenize(tokenizer, prompt, words): | ||
| 73 | + """Tokenize prompt+words and return input_ids, attention_mask, words_mask. | ||
| 74 | + | ||
| 75 | + Mirrors GLiNER's tokenize_inputs + prepare_word_mask. | ||
| 76 | + """ | ||
| 77 | + tokens = prompt + [w for w, _, _ in words] | ||
| 78 | + enc = tokenizer(tokens, is_split_into_words=True, add_special_tokens=True, | ||
| 79 | + truncation=True, max_length=SEQ_LEN) | ||
| 80 | + input_ids = np.asarray(enc["input_ids"], dtype=np.int32) | ||
| 81 | + attention_mask = np.asarray(enc["attention_mask"], dtype=np.int32) | ||
| 82 | + word_ids = enc.word_ids() | ||
| 83 | + | ||
| 84 | + words_mask = np.zeros_like(input_ids, dtype=np.int32) | ||
| 85 | + prev_wid = None | ||
| 86 | + seen_words = 0 | ||
| 87 | + skip_n = len(prompt) | ||
| 88 | + for i, wid in enumerate(word_ids): | ||
| 89 | + if wid is None: | ||
| 90 | + prev_wid = wid | ||
| 91 | + continue | ||
| 92 | + if wid != prev_wid: | ||
| 93 | + seen_words += 1 | ||
| 94 | + prev_wid = wid | ||
| 95 | + if seen_words > skip_n and words_mask[i] == 0 and (i == 0 or word_ids[i - 1] != wid): | ||
| 96 | + words_mask[i] = seen_words - skip_n | ||
| 97 | + return input_ids, attention_mask, words_mask | ||
| 98 | + | ||
| 99 | + | ||
| 100 | +def _prepare_span_idx(): | ||
| 101 | + """Pre-compute the fixed (NUM_WORDS * MAX_WIDTH, 2) span index grid.""" | ||
| 102 | + starts = np.arange(NUM_WORDS, dtype=np.int32).reshape(-1, 1) | ||
| 103 | + offsets = np.arange(MAX_WIDTH, dtype=np.int32).reshape(1, -1) | ||
| 104 | + grid = np.stack([ | ||
| 105 | + np.broadcast_to(starts, (NUM_WORDS, MAX_WIDTH)), | ||
| 106 | + np.broadcast_to(starts + offsets, (NUM_WORDS, MAX_WIDTH)), | ||
| 107 | + ], axis=-1) | ||
| 108 | + return grid.reshape(-1, 2) | ||
| 109 | + | ||
| 110 | + | ||
| 111 | +def _prepare_inputs(text, labels, tokenizer, ent_token, sep_token): | ||
| 112 | + """Build the 6 MindSpore Lite inputs for a single (text, labels) pair.""" | ||
| 113 | + all_words = _split_words(text) | ||
| 114 | + prompt = _build_prompt(labels, ent_token, sep_token) | ||
| 115 | + | ||
| 116 | + # Truncate body words to fit both SEQ_LEN (tokens) and NUM_WORDS (word slots). | ||
| 117 | + words = all_words[:NUM_WORDS] | ||
| 118 | + while words: | ||
| 119 | + input_ids, attention_mask, words_mask = _tokenize(tokenizer, prompt, words) | ||
| 120 | + if input_ids.shape[0] <= SEQ_LEN: | ||
| 121 | + break | ||
| 122 | + words = words[:-1] | ||
| 123 | + num_body_words = len(words) | ||
| 124 | + | ||
| 125 | + # Pad seq dim up to SEQ_LEN. | ||
| 126 | + pad_len = SEQ_LEN - input_ids.shape[0] | ||
| 127 | + if pad_len > 0: | ||
| 128 | + input_ids = np.concatenate([input_ids, np.zeros(pad_len, dtype=np.int32)]) | ||
| 129 | + attention_mask = np.concatenate([attention_mask, np.zeros(pad_len, dtype=np.int32)]) | ||
| 130 | + words_mask = np.concatenate([words_mask, np.zeros(pad_len, dtype=np.int32)]) | ||
| 131 | + | ||
| 132 | + text_lengths = np.asarray([[num_body_words]], dtype=np.int32) | ||
| 133 | + | ||
| 134 | + span_idx = _prepare_span_idx() | ||
| 135 | + span_mask = (span_idx[:, 1] < num_body_words).reshape(1, -1).astype(bool) | ||
| 136 | + | ||
| 137 | + feeds = [ | ||
| 138 | + input_ids.reshape(1, -1), | ||
| 139 | + attention_mask.reshape(1, -1), | ||
| 140 | + words_mask.reshape(1, -1), | ||
| 141 | + text_lengths, | ||
| 142 | + span_idx.reshape(1, -1, 2), | ||
| 143 | + span_mask, | ||
| 144 | + ] | ||
| 145 | + return feeds, words, num_body_words | ||
| 146 | + | ||
| 147 | + | ||
| 148 | +def _greedy_search(spans, flat_ner): | ||
| 149 | + """Greedy non-overlap selection. spans: list of (start, end, label, score).""" | ||
| 150 | + spans_sorted = sorted(spans, key=lambda s: s[3], reverse=True) | ||
| 151 | + picked = [] | ||
| 152 | + for span in spans_sorted: | ||
| 153 | + start, end, _, _ = span | ||
| 154 | + if flat_ner: | ||
| 155 | + conflict = any(not (end < p_start or start > p_end) | ||
| 156 | + for p_start, p_end, _, _ in picked) | ||
| 157 | + if conflict: | ||
| 158 | + continue | ||
| 159 | + picked.append(span) | ||
| 160 | + return picked | ||
| 161 | + | ||
| 162 | + | ||
| 163 | +def _decode(logits, num_body_words, labels, threshold, flat_ner): | ||
| 164 | + """Decode (1, NUM_WORDS, MAX_WIDTH, NUM_CLASSES) logits to char spans.""" | ||
| 165 | + probs = 1.0 / (1.0 + np.exp(-logits[0])) | ||
| 166 | + candidates = [] | ||
| 167 | + for s in range(num_body_words): | ||
| 168 | + for k in range(MAX_WIDTH): | ||
| 169 | + if s + k >= num_body_words: | ||
| 170 | + continue | ||
| 171 | + for c, label in enumerate(labels): | ||
| 172 | + score = float(probs[s, k, c]) | ||
| 173 | + if score >= threshold: | ||
| 174 | + candidates.append((s, s + k, label, score)) | ||
| 175 | + return _greedy_search(candidates, flat_ner) | ||
| 176 | + | ||
| 177 | + | ||
| 178 | +def _map_to_chars(spans, words): | ||
| 179 | + """Map word-index spans (start, end, label, score) -> char dict list.""" | ||
| 180 | + out = [] | ||
| 181 | + for s, e, label, score in spans: | ||
| 182 | + char_start = words[s][1] | ||
| 183 | + char_end = words[e][2] | ||
| 184 | + out.append({ | ||
| 185 | + "text": words[s][0] if s == e else " ".join(w[0] for w in words[s:e + 1]), | ||
| 186 | + "label": label, | ||
| 187 | + "score": score, | ||
| 188 | + "start": char_start, | ||
| 189 | + "end": char_end, | ||
| 190 | + }) | ||
| 191 | + return out | ||
| 192 | + | ||
| 193 | + | ||
| 194 | +def load_config(model_dir): | ||
| 195 | + """Load ent_token, sep_token, max_width from gliner_config.json.""" | ||
| 196 | + cfg_path = Path(model_dir) / "gliner_config.json" | ||
| 197 | + with open(cfg_path, "r", encoding="utf-8") as f: | ||
| 198 | + cfg = json.load(f) | ||
| 199 | + return cfg | ||
| 200 | + | ||
| 201 | + | ||
| 202 | +def parse_args(): | ||
| 203 | + """Parse command-line arguments.""" | ||
| 204 | + parser = argparse.ArgumentParser(description="gliner_large-v2.5 MindSpore Lite inference") | ||
| 205 | + parser.add_argument("--model-dir", default="gliner_large-v2.5", | ||
| 206 | + help="Path to the original gliner checkpoint (for tokenizer + config)") | ||
| 207 | + parser.add_argument("--mindir-path", default="./onnx/model.mindir", | ||
| 208 | + help="Path to the converted MindIR model") | ||
| 209 | + parser.add_argument("--text", default=None, help="Input text (overrides --text-file)") | ||
| 210 | + parser.add_argument("--text-file", default=None, help="File with one text per line") | ||
| 211 | + parser.add_argument("--labels", default=",".join(DEFAULT_LABELS), | ||
| 212 | + help="Comma-separated entity labels (must match export-time labels)") | ||
| 213 | + parser.add_argument("--threshold", type=float, default=0.5, | ||
| 214 | + help="Confidence threshold (sigmoid probability)") | ||
| 215 | + parser.add_argument("--flat-ner", action="store_true", default=True, | ||
| 216 | + help="Disallow overlapping spans (default True)") | ||
| 217 | + parser.add_argument("--device-id", type=int, default=0, help="Ascend device id") | ||
| 218 | + return parser.parse_args() | ||
| 219 | + | ||
| 220 | + | ||
| 221 | +def main(): | ||
| 222 | + """Run gliner_large-v2.5 MindSpore Lite inference end-to-end.""" | ||
| 223 | + args = parse_args() | ||
| 224 | + cfg = load_config(args.model_dir) | ||
| 225 | + ent_token = cfg["ent_token"] | ||
| 226 | + sep_token = cfg["sep_token"] | ||
| 227 | + labels = [s.strip() for s in args.labels.split(",") if s.strip()] | ||
| 228 | + if len(labels) != NUM_CLASSES: | ||
| 229 | + raise ValueError( | ||
| 230 | + f"Static MindIR has {NUM_CLASSES} classes baked in, but got {len(labels)} labels: {labels}. " | ||
| 231 | + f"Re-export ONNX with --labels and re-convert to change the label set." | ||
| 232 | + ) | ||
| 233 | + print(f"[infer] config: ent='{ent_token}', sep='{sep_token}', labels={labels}", flush=True) | ||
| 234 | + | ||
| 235 | + print(f"[infer] loading tokenizer from {args.model_dir}", flush=True) | ||
| 236 | + tokenizer = AutoTokenizer.from_pretrained(args.model_dir, use_fast=True) | ||
| 237 | + | ||
| 238 | + print(f"[infer] loading MindIR from {args.mindir_path}", flush=True) | ||
| 239 | + context = mslite.Context() | ||
| 240 | + context.target = ["ascend"] | ||
| 241 | + context.ascend.device_id = args.device_id | ||
| 242 | + model = mslite.Model() | ||
| 243 | + model.build_from_file(args.mindir_path, mslite.ModelType.MINDIR, context) | ||
| 244 | + | ||
| 245 | + if args.text is not None: | ||
| 246 | + texts = [args.text] | ||
| 247 | + elif args.text_file is not None: | ||
| 248 | + with open(args.text_file, "r", encoding="utf-8") as f: | ||
| 249 | + texts = [line.rstrip("\n") for line in f if line.strip()] | ||
| 250 | + else: | ||
| 251 | + texts = [ | ||
| 252 | + "Cristiano Ronaldo dos Santos Aveiro plays for Al-Nassr FC and captains Portugal.", | ||
| 253 | + "Linus Torvalds created Linux in 1991 while at the University of Helsinki.", | ||
| 254 | + "The Eiffel Tower is located in Paris, France and was built in 1889.", | ||
| 255 | + ] | ||
| 256 | + | ||
| 257 | + for text in texts: | ||
| 258 | + feeds, words, num_body_words = _prepare_inputs( | ||
| 259 | + text, labels, tokenizer, ent_token, sep_token) | ||
| 260 | + outputs = model.predict(feeds) | ||
| 261 | + logits = outputs[0].get_data_to_numpy() | ||
| 262 | + spans = _decode(logits, num_body_words, labels, args.threshold, args.flat_ner) | ||
这个decode的逻辑需要入图吗?这个逻辑是针对这个场景定制的还是所有的场景都会执行的 ![]() ![]() | |||
| 263 | + ents = _map_to_chars(spans, words) | ||
| 264 | + print(f"\n[infer] text: {text}", flush=True) | ||
| 265 | + print(f"[infer] words ({num_body_words}): {[w[0] for w in words]}", flush=True) | ||
| 266 | + print(f"[infer] seq_len: {feeds[0].shape[1]}, logits shape: {logits.shape}", flush=True) | ||
| 267 | + for ent in ents: | ||
| 268 | + print(f" - {ent['text']!r} [{ent['label']}] score={ent['score']:.4f} " | ||
| 269 | + f"chars=({ent['start']}, {ent['end']})", flush=True) | ||
| 270 | + | ||
| 271 | + | ||
| 272 | +if __name__ == "__main__": | ||
| 273 | + main() | ||


🟡 Medium Priority
changed line:
infer_gliner_large-v2.5_mslite.pyL118-124while words:循环体affected behavior: 当
text为空字符串时,_split_words("")返回空列表[],words = all_words[:NUM_WORDS]也为[],导致 while 循环条件为 False 直接跳过,input_ids、attention_mask、words_mask三个变量从未被赋值。随后 L127pad_len = SEQ_LEN - input_ids.shape[0]访问未定义变量,抛出NameError。failure mode: 用户通过
--text ""传入空字符串时,脚本直接崩溃,错误信息不友好(NameError: name 'input_ids' is not defined),而非给出明确的参数校验提示。suggested fix: 在
_prepare_inputs开头的all_words = _split_words(text)之后增加空文本检测,若len(all_words) == 0则抛出ValueError("text must contain at least one word"),或在main()中遍历 texts 时对空文本 skip + warning。建议:在
_prepare_inputs中all_words = _split_words(text)之后添加空文本检测:if not all_words: raise ValueError("text must contain at least one word")