已关闭
[Bug]: 310P上使用tei 26.0.0镜像运行qwen3系列embed和reranker报错 #130
zhouxiao999创建于  8月12日关闭于  8月14日
zhouxiao999
8月12日 创建

环境信息

例如:
- 操作系统:Ubuntu 20.04.5 LTS
- 昇腾硬件信息:arm机器,300IDUO卡
- 昇腾驱动:一开始Ascend HDK 25.2.0,后来升级到最新版  26.1.0 问题依旧
- 安装的对应软件版本:docker pull --platform=arm64 swr.cn-south-1.myhuaweicloud.com/ascendhub/mis-tei:26.0.0-310p-ubuntu22.04-py3.11和docker pull --platform=arm64 swr.cn-south-1.myhuaweicloud.com/ascendhub/mis-tei:26.0.0-310p-openeuler24.03-py3.11都报错

🐛 问题描述

使用官方tei 26.0.0在310p机器上运行Qwen3-Embedding-0.6B和Qwen3-Reranker-0.6B都报错RotaryMul1算子异常
启动镜像的命令:

services:
  tei:
    image: swr.cn-south-1.myhuaweicloud.com/ascendhub/mis-tei:26.0.0-310p-ubuntu22.04-py3.11
    container_name: tei26-qwen3-embedding-debug
    user: root
    network_mode: host
    environment:
      HOME: /home/HwHiAiUser
      MODEL_DIR: /home/HwHiAiUser/model
    devices:
      - /dev/davinci4
      - /dev/davinci_manager
      - /dev/devmm_svm
      - /dev/hisi_hdc
    volumes:
      - /usr/local/dcmi:/usr/local/dcmi
      - /usr/local/bin/npu-smi:/usr/local/bin/npu-smi
      - /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64
      - /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info
      - /etc/ascend_install.info:/etc/ascend_install.info
      - /home/data/nlp_openmodel:/home/HwHiAiUser/model:ro
    command:
      - --model-id
      - /home/HwHiAiUser/model/Qwen3-Embedding-0.6B
      - --hostname
      - 0.0.0.0
      - --port
      - "8095"    
      - --max-batch-tokens
      - "32768"
    restart: always

报错信息如下:

(base) root@atlas-duo:/home/data/zx/tei/tei26# docker compose up
[+] Running 1/0
 ✔ Container tei26-qwen3-embedding-debug  Created                                                                                       0.0s 
Attaching to tei26-qwen3-embedding-debug
tei26-qwen3-embedding-debug  | Using local model path: /home/HwHiAiUser/model/Qwen3-Embedding-0.6B
tei26-qwen3-embedding-debug  | Model '/home/HwHiAiUser/model/Qwen3-Embedding-0.6B' exists.
tei26-qwen3-embedding-debug  | Starting TEI service with model 'Qwen3-Embedding-0.6B'...
tei26-qwen3-embedding-debug  | Additional arguments:  --hostname 0.0.0.0 --port 8095 --max-batch-tokens 32768
tei26-qwen3-embedding-debug  | Executing: text-embeddings-router --model-id Qwen3-Embedding-0.6B  --hostname 0.0.0.0 --port 8095 --max-batch-tokens 32768
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:18.946106Z  INFO text_embeddings_router: router/src/main.rs:216: Args { model_id: "Qwe**-*********-0.6B", revision: None, tokenization_workers: None, dtype: None, served_model_name: None, pooling: None, max_concurrent_requests: 512, max_batch_tokens: 32768, max_batch_requests: None, max_client_batch_size: 32, auto_truncate: true, default_prompt_name: None, default_prompt: None, dense_path: None, hf_api_token: None, hf_token: None, hostname: "0.0.0.0", port: 8095, uds_path: "/tmp/text-embeddings-inference-server", huggingface_hub_cache: None, payload_limit: 2000000, api_key: None, json_output: false, disable_spans: false, otlp_endpoint: None, otlp_service_name: "text-embeddings-inference.server", prometheus_port: 9000, cors_allow_origin: None }
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:19.508832Z  WARN text_embeddings_router: router/src/lib.rs:217: Could not find a Sentence Transformers config
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:19.508863Z  INFO text_embeddings_router: router/src/lib.rs:235: Maximum number of tokens per request: 32768
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:19.509212Z  INFO text_embeddings_core::tokenization: core/src/tokenization.rs:38: Starting 64 tokenization workers
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:19.509384Z  INFO text_embeddings_router: router/src/lib.rs:285: Starting model backend
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:19.509672Z  INFO text_embeddings_backend_python::management: backends/python/src/management.rs:68: Starting Python backend
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:29.521217Z  INFO text_embeddings_backend_python::management: backends/python/src/management.rs:132: Waiting for Python backend to be ready...
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:39.525673Z  INFO text_embeddings_backend_python::management: backends/python/src/management.rs:132: Waiting for Python backend to be ready...
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:42.420062Z  INFO python-backend: text_embeddings_backend_python::logging: backends/python/src/logging.rs:37: backend device: npu
tei26-qwen3-embedding-debug  | 
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:44.471825Z  INFO python-backend: text_embeddings_backend_python::logging: backends/python/src/logging.rs:37: Server started at unix:///tmp/text-embeddings-inference-server
tei26-qwen3-embedding-debug  | 
tei26-qwen3-embedding-debug  | 2026-08-12T08:47:44.473181Z  INFO text_embeddings_backend_python::management: backends/python/src/management.rs:129: Python backend ready in 24.956044239s
tei26-qwen3-embedding-debug  | 2026-08-12T08:49:08.823076Z ERROR health:embed:embed: backend_grpc_client: backends/grpc-client/src/lib.rs:25: Server error: The Inner error is reported as above. The process exits for this inner error, and the current working operator name is RotaryMul.
tei26-qwen3-embedding-debug  | Since the operator is called asynchronously, the stacktrace may be inaccurate. If you want to get the accurate stacktrace, please set the environment variable ASCEND_LAUNCH_BLOCKING=1.
tei26-qwen3-embedding-debug  | Note: ASCEND_LAUNCH_BLOCKING=1 will force ops to run in synchronous mode, resulting in performance degradation. Please unset ASCEND_LAUNCH_BLOCKING in time after debugging.
tei26-qwen3-embedding-debug  | [ERROR] 2026-08-12-08:49:08 (PID:315, Device:0, RankID:-1) ERR00100 PTA call acl api failed.
tei26-qwen3-embedding-debug  | [PID: 315] 2026-08-12-08:49:08.807.416 Compilation_Error(E40021): Failed to compile Op RotaryMul1. oppath is /usr/local/Ascend/cann-9.0.0/opp/built-in/op_impl/ai_core/tbe/impl/ops_legacy/dynamic/rotary_mul.py Pre-compile failed with errormsg/stack:  and optype is RotaryMul.
tei26-qwen3-embedding-debug  |         Solution: See the host log for details, and then check the Python stack where the error log is reported.
tei26-qwen3-embedding-debug  |         TraceBack (most recent call last):
tei26-qwen3-embedding-debug  |         Pre-compile op[RotaryMul1] failed, oppath[/usr/local/Ascend/cann-9.0.0/opp/built-in/op_impl/ai_core/tbe/impl/ops_legacy/dynamic/rotary_mul.py], optype[RotaryMul], taskID[2]. Please check op's compilation error message.[FUNC:ReportBuildErrMessage][FILE:fusion_manager.cc][LINE:374]
tei26-qwen3-embedding-debug  |         [SubGraphOpt][Compile][ProcFailedCompTask] Thread[281459767832864] failed to recompile single op[RotaryMul1][FUNC:ProcessAllFailedCompileTasks][FILE:tbe_op_store_adapter.cc][LINE:1169]
tei26-qwen3-embedding-debug  |         [SubGraphOpt][Compile][ParalCompOp] Thread[281459767832864] failed when processing the task that had failed to compile.[FUNC:ParallelCompileOp][FILE:tbe_op_store_adapter.cc][LINE:1216]
tei26-qwen3-embedding-debug  |         [SubGraphOpt][Compile][CompOpOnly] CompileOp failed.[FUNC:CompileOpOnly][FILE:op_compiler.cc][LINE:1175]
tei26-qwen3-embedding-debug  |         [GraphOpt][FusedGraph][RunCompile] Failed to compile graph with compiler Normal mode Op Compiler[FUNC:SubGraphCompile][FILE:fe_graph_optimizer.cc][LINE:1452]
tei26-qwen3-embedding-debug  |         Call OptimizeFusedGraph failed, ret:4294967295, engine_name:AIcoreEngine, graph_name:partition0_rank1_new_sub_graph3[FUNC:OptimizeSubGraph][FILE:graph_optimize.cc][LINE:124]
tei26-qwen3-embedding-debug  |         subgraph 0 optimize failed[FUNC:OptimizeSubGraphWithMultiThreads][FILE:graph_manager.cc][LINE:1032]
tei26-qwen3-embedding-debug  |         Assert ((DoSubgraphPartitionWithMode(graph_node, compute_graph, session_id, EnginePartitioner::Mode::kAtomicEnginePartitioning, "AtomicEngine")) == ge::SUCCESS) failed[FUNC:OptimizeSubgraph][FILE:graph_manager.cc][LINE:3898]
tei26-qwen3-embedding-debug  |         build graph failed, graph id:0, ret:1343225857[FUNC:BuildModelWithGraphId][FILE:ge_generator.cc][LINE:1603]
tei26-qwen3-embedding-debug  |         [Build][SingleOpModel]call ge interface generator.BuildSingleOpModel failed. ge result = 1343225857[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:148]
tei26-qwen3-embedding-debug  |         [Build][Op]Fail to build op model[FUNC:ReportInnerError][FILE:log_inner.cpp][LINE:132]
tei26-qwen3-embedding-debug  |         build op model failed, result = 500002[FUNC:ReportInnerError][FILE:log_inner.cpp][LINE:132]
tei26-qwen3-embedding-debug  | 
tei26-qwen3-embedding-debug  | 2026-08-12T08:49:08.963229Z  INFO text_embeddings_backend_python::management: backends/python/src/management.rs:146: Python backend process terminated
tei26-qwen3-embedding-debug  | Error: Model backend is not healthy
tei26-qwen3-embedding-debug  | 
tei26-qwen3-embedding-debug  | Caused by:
tei26-qwen3-embedding-debug  |     Server error: The Inner error is reported as above. The process exits for this inner error, and the current working operator name is RotaryMul.
tei26-qwen3-embedding-debug  |     Since the operator is called asynchronously, the stacktrace may be inaccurate. If you want to get the accurate stacktrace, please set the environment variable ASCEND_LAUNCH_BLOCKING=1.
tei26-qwen3-embedding-debug  |     Note: ASCEND_LAUNCH_BLOCKING=1 will force ops to run in synchronous mode, resulting in performance degradation. Please unset ASCEND_LAUNCH_BLOCKING in time after debugging.
tei26-qwen3-embedding-debug  |     [ERROR] 2026-08-12-08:49:08 (PID:315, Device:0, RankID:-1) ERR00100 PTA call acl api failed.
tei26-qwen3-embedding-debug  |     [PID: 315] 2026-08-12-08:49:08.807.416 Compilation_Error(E40021): Failed to compile Op RotaryMul1. oppath is /usr/local/Ascend/cann-9.0.0/opp/built-in/op_impl/ai_core/tbe/impl/ops_legacy/dynamic/rotary_mul.py Pre-compile failed with errormsg/stack:  and optype is RotaryMul.
tei26-qwen3-embedding-debug  |             Solution: See the host log for details, and then check the Python stack where the error log is reported.
tei26-qwen3-embedding-debug  |             TraceBack (most recent call last):
tei26-qwen3-embedding-debug  |             Pre-compile op[RotaryMul1] failed, oppath[/usr/local/Ascend/cann-9.0.0/opp/built-in/op_impl/ai_core/tbe/impl/ops_legacy/dynamic/rotary_mul.py], optype[RotaryMul], taskID[2]. Please check op's compilation error message.[FUNC:ReportBuildErrMessage][FILE:fusion_manager.cc][LINE:374]
tei26-qwen3-embedding-debug  |             [SubGraphOpt][Compile][ProcFailedCompTask] Thread[281459767832864] failed to recompile single op[RotaryMul1][FUNC:ProcessAllFailedCompileTasks][FILE:tbe_op_store_adapter.cc][LINE:1169]
tei26-qwen3-embedding-debug  |             [SubGraphOpt][Compile][ParalCompOp] Thread[281459767832864] failed when processing the task that had failed to compile.[FUNC:ParallelCompileOp][FILE:tbe_op_store_adapter.cc][LINE:1216]
tei26-qwen3-embedding-debug  |             [SubGraphOpt][Compile][CompOpOnly] CompileOp failed.[FUNC:CompileOpOnly][FILE:op_compiler.cc][LINE:1175]
tei26-qwen3-embedding-debug  |             [GraphOpt][FusedGraph][RunCompile] Failed to compile graph with compiler Normal mode Op Compiler[FUNC:SubGraphCompile][FILE:fe_graph_optimizer.cc][LINE:1452]
tei26-qwen3-embedding-debug  |             Call OptimizeFusedGraph failed, ret:4294967295, engine_name:AIcoreEngine, graph_name:partition0_rank1_new_sub_graph3[FUNC:OptimizeSubGraph][FILE:graph_optimize.cc][LINE:124]
tei26-qwen3-embedding-debug  |             subgraph 0 optimize failed[FUNC:OptimizeSubGraphWithMultiThreads][FILE:graph_manager.cc][LINE:1032]
tei26-qwen3-embedding-debug  |             Assert ((DoSubgraphPartitionWithMode(graph_node, compute_graph, session_id, EnginePartitioner::Mode::kAtomicEnginePartitioning, "AtomicEngine")) == ge::SUCCESS) failed[FUNC:OptimizeSubgraph][FILE:graph_manager.cc][LINE:3898]
tei26-qwen3-embedding-debug  |             build graph failed, graph id:0, ret:1343225857[FUNC:BuildModelWithGraphId][FILE:ge_generator.cc][LINE:1603]
tei26-qwen3-embedding-debug  |             [Build][SingleOpModel]call ge interface generator.BuildSingleOpModel failed. ge result = 1343225857[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:148]
tei26-qwen3-embedding-debug  |             [Build][Op]Fail to build op model[FUNC:ReportInnerError][FILE:log_inner.cpp][LINE:132]
tei26-qwen3-embedding-debug  |             build op model failed, result = 500002[FUNC:ReportInnerError][FILE:log_inner.cpp][LINE:132]
tei26-qwen3-embedding-debug  |     
tei26-qwen3-embedding-debug exited with code 0

欢迎加入社区,感谢您对社区的贡献 🎉!

likedislike
xiangjie10成员
8月12日 评论:

👋 您好,感谢向 text-embeddings-inference 提交 Issue!
🎉 我们已收到您的反馈,感谢你对开源社区的支持!

📅 处理时效 维护团队将在工作日 24 小时内查看并回复您的问题。
🔍 自助排查(推荐优先查看) 在等待回复期间,您可以先查阅仓库README以及历史 Issue 中相似问题的解决方案,多数问题可快速解决。
💡 为了更快定位问题,请您确保 Issue 包含:

  • 清晰的问题描述
  • 可复现的操作步骤
  • 相关日志、截图或环境信息
    我们会尽快跟进,感谢您的理解与配合!
likedislike
xiangjie10成员
8月12日 评论:

/label add triaged

likedislike
ascend-robotascend-robot成员
8月12日 添加了label:bug
ascend-robotascend-robot成员
8月12日 添加了label:triaged
Zzhouxiao999
8月12日 修改了issue 的描述
wangting
wangting成员
8月13日 评论:

感谢反馈,该问题我们正在排查。目前可先回退使用之前版本的镜像作为临时方案,后续有修复进展会及时同步。26.0.0之前的版本启动命令,可以参考:https://gitcode.com/wangyongjun/text-embeddings-inference/blob/add_26.1.0_version/docker/OVERVIEW.zh.md

likedislike
wangting
wangting成员
8月14日 评论:

您好,我们在310p上面复测多次发现功能都是正常的,没有出现你们说的报错,这是我们使用的启动命令
export MODEL_DIR=/data/models
docker run
-u root
--privileged
--net host
--name tei_container
--device /dev/davinci0
--device /dev/davinci_manager
--device /dev/devmm_svm
--device /dev/hisi_hdc
-v /usr/local/dcmi:/usr/local/dcmi
-v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info
-v /etc/ascend_install.info:/etc/ascend_install.info
-e MODEL_DIR=$MODEL_DIR
-v MODELDIR:MODEL_DIR:MODEL_DIR
-itd swr.cn-south-1.myhuaweicloud.com/ascendhub/mis-tei:26.0.0-310p-ubuntu22.04-py3.11 --model-id $MODEL_DIR/Qwen3-Embedding-0.6B --hostname 51.38.4.5 --port 5005

likedislike
wangtingwangting成员
8月14日 issue状态由 TODO 改变为 DONE
wangtingwangting成员
8月14日 关闭了 issue
zhouxiao999
8月14日 评论:

感谢,应该是卡的冲突或者docker compose文件有点小问题导致的,参考楼上对比修改后能正常运行了

services:
  tei:
    image: swr.cn-south-1.myhuaweicloud.com/ascendhub/mis-tei:26.0.0-310p-ubuntu22.04-py3.11
    container_name: tei26-qwen3-embedding
    user: root
    # privileged: true
    network_mode: host
    environment:
      MODEL_DIR: /home/HwHiAiUser/model
    devices:
      - /dev/davinci2
      - /dev/davinci_manager
      - /dev/devmm_svm
      - /dev/hisi_hdc
    volumes:
      - /usr/local/dcmi:/usr/local/dcmi
      - /usr/local/bin/npu-smi:/usr/local/bin/npu-smi
      - /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64
      - /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info
      - /etc/ascend_install.info:/etc/ascend_install.info
      - /home/data/nlp_openmodel:/home/HwHiAiUser/model
    command:
      - --model-id
      - /home/HwHiAiUser/model/Qwen3-Embedding-0.6B
      - --hostname
      - 0.0.0.0
      - --port
      - "5006"
services:
  tei:
    image: swr.cn-south-1.myhuaweicloud.com/ascendhub/mis-tei:26.0.0-310p-ubuntu22.04-py3.11
    container_name: tei26-qwen3-reranker
    user: root
    network_mode: host
    environment:
      MODEL_DIR: /home/HwHiAiUser/model
      IS_RERANK: "1"
      DEFAULT_PROMPT: |-
        <Instruct>: Given a web search query, retrieve relevant passages that answer the query
        <Query>: <s1>
        <Document>: <s2>
      # LOG_LEVEL: "info,text_embeddings_core::queue=warn,text_embeddings_router::http::server=warn"
    devices:
      - /dev/davinci3
      - /dev/davinci_manager
      - /dev/devmm_svm
      - /dev/hisi_hdc
    volumes:
      - /usr/local/dcmi:/usr/local/dcmi
      - /usr/local/bin/npu-smi:/usr/local/bin/npu-smi
      - /usr/local/Ascend/driver/lib64:/usr/local/Ascend/driver/lib64
      - /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info
      - /etc/ascend_install.info:/etc/ascend_install.info
      - /home/data/nlp_openmodel:/home/HwHiAiUser/model
      - ./sentence_bert_config.json:/home/HwHiAiUser/model/Qwen3-Reranker-0.6B/sentence_bert_config.json:ro
    command:
      - --model-id
      - /home/HwHiAiUser/model/Qwen3-Reranker-0.6B
      - --hostname
      - 0.0.0.0
      - --port
      - "5007"
      - --dtype
      - float16
      - --max-batch-tokens
      - "65536"
    restart: always

likedislike