已合并
feat(docker): 评测镜像跨架构 + 9.1.0 + 950PR 支持 #239
feat(docker): 评测镜像跨架构 + 9.1.0 + 950PR 支持 #239
已合并
Xinxian Chen创建于 7月31日
共 10 个文件变更+355-88
@@ -8,6 +8,18 @@ base/ (cann-toolkit-base) 环境底座: CANN toolkit + torch/torch_npu
8dev/ (cann-bench:cann9.0.0-*) AscendHub 全量 CANN 的交互/CI 调试镜像 (独立血统)8dev/ (cann-bench:cann9.0.0-*) AscendHub 全量 CANN 的交互/CI 调试镜像 (独立血统)
9```9```
10 10 
11+**base 就是 common。** 它到 toolkit 为止不装 ops,唯一的变量是 CPU 架构(`ARCH`,写进
12+`ENV CANN_ARCH` 供下游继承)。**芯片(SoC)的分叉正好从 `ops.run` 开始**,所以整条分叉都在
13+`eval/` 里,而且只是三个 ARG 值(`NPU_ARCH` / `OPS_PKG` / `OPS_MODE`)—— a2 与 a5 的镜像**零
14+结构差异**,不需要两份 Dockerfile。
15+ 
16+| 目标 | 建在哪 | `ARCH` | `NPU_ARCH` | `OPS_MODE` |
17+|---|---|---|---|---|
18+| A2 (910B2) | aarch64 host | `aarch64` | `ascend910b` | `none` |
19+| A5 (950PR) | x86_64 host | `x86_64` | `ascend950` | **`refonly`**(`none` 在 950 上跑不了,见 eval README) |
20+ 
21+原生构建,没有交叉编译 —— base 和 eval 必须在目标架构的机器上建,不符会在 build 期直接失败。
22+ 
11| | 干什么 | 什么时候用 |23| | 干什么 | 什么时候用 |
12|---|---|---|24|---|---|---|
13| [`eval/`](eval/) | **`docker run <image> [源码目录] [选项]` 直接产出评测报告** | 评一个提交;CI 打分;任何要求"这个分数出自哪个 benchmark 版本"可回答的场景 |25| [`eval/`](eval/) | **`docker run <image> [源码目录] [选项]` 直接产出评测报告** | 评一个提交;CI 打分;任何要求"这个分数出自哪个 benchmark 版本"可回答的场景 |
@@ -27,6 +27,12 @@ ARG UV_PYTHON_INSTALL_MIRROR=
27ARG PYPI_MIRROR=27ARG PYPI_MIRROR=
28ARG TORCH_MIRROR=28ARG TORCH_MIRROR=
29ARG CANN_VERSION=9.0.129ARG CANN_VERSION=9.0.1
30+# CANN_VERSION alone is enough to move versions -- the toolkit .run and its <arch> variants follow one
31+# naming scheme on ascend-repo, verified present for 9.1.0 in both architectures. CANN_TOOLKIT_URL is
32+# the escape hatch for the releases that do not: some land on the ascend-cann-open bucket instead
33+# (ascend-cann-open.obs.cn-north-4...), and some are published as a combined `Ascend-cann_<v>_linux-
34+# <arch>.run` rather than `Ascend-cann-toolkit_...`. Empty = derive from CANN_VERSION + ARCH.
35+ARG CANN_TOOLKIT_URL=
30 36 
31FROM ${UV_IMAGE} AS uvbin37FROM ${UV_IMAGE} AS uvbin
32FROM ${BASE_OS}38FROM ${BASE_OS}
@@ -37,11 +43,19 @@ ARG UV_PYTHON_INSTALL_MIRROR
37ARG PYPI_MIRROR43ARG PYPI_MIRROR
38ARG TORCH_MIRROR44ARG TORCH_MIRROR
39ARG CANN_VERSION45ARG CANN_VERSION
46+ARG CANN_TOOLKIT_URL
40 47 
41-# WHY aarch64 is hardcoded (not an ARG): the image is aarch64-only by construction -- torch/torch_npu48+# ARCH selects the CPU architecture of every arch-shaped path: the toolkit .run, the toolkit's
42-# wheels, --platform, and every tested path are aarch64 (Ascend hosts are Kunpeng/ARM). A parametric49+# <arch>-linux/ tree, and the wheels uv picks. Build it NATIVELY on a host of that architecture --
43-# ARCH knob we never build on x86_64 is a false promise (silent FHS-path breakage in the JSON ENTRYPOINT50+# this is not a cross-compile knob, and there is no --platform anywhere.
44-# that can't expand build-args); an x86_64 image, if ever needed, is a separate deliberately-tested variant.51+#
52+# It was hardcoded aarch64 until x86_64 was actually exercised (Ascend hosts are usually Kunpeng/ARM,
53+# and the old comment rightly refused a knob nobody built). Both are now validated end to end:
54+# aarch64/910B2 and x86_64/950PR each install the toolkit, `uv sync --frozen` off the same lock (it
55+# already carries both architectures' wheels, torch_npu included), stub libhccl, build cann_bench_utils,
56+# and see their card. The other half of that old objection -- "a JSON ENTRYPOINT can't expand
57+# build-args" -- is answered below by baking the arch into a generated env script instead.
58+ARG ARCH=aarch64
45 59 
46ENV DEBIAN_FRONTEND=noninteractive60ENV DEBIAN_FRONTEND=noninteractive
47SHELL ["/bin/bash", "-c"]61SHELL ["/bin/bash", "-c"]
@@ -77,24 +91,60 @@ ENV PATH=/opt/venv/bin:${PATH}
77 91 
78# Toolkit ONLY -- the ops/nnal .run the upstream cann-container-image also fetches are omitted.92# Toolkit ONLY -- the ops/nnal .run the upstream cann-container-image also fetches are omitted.
79RUN cd /tmp \93RUN cd /tmp \
80- && URL="https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/CANN/CANN%20${CANN_VERSION}/Ascend-cann-toolkit_${CANN_VERSION}_linux-aarch64.run" \94+ && URL="${CANN_TOOLKIT_URL:-https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/CANN/CANN%20${CANN_VERSION}/Ascend-cann-toolkit_${CANN_VERSION}_linux-${ARCH}.run}" \
95+ && echo "toolkit: ${URL}" \
81 && wget --quiet --header="Referer: https://www.hiascend.com/" -O toolkit.run "${URL}" \96 && wget --quiet --header="Referer: https://www.hiascend.com/" -O toolkit.run "${URL}" \
82 && chmod +x toolkit.run && ./toolkit.run --quiet --install --install-for-all && rm -f toolkit.run97 && chmod +x toolkit.run && ./toolkit.run --quiet --install --install-for-all && rm -f toolkit.run
83 98 
84# WHY stub: torch_npu _C.so has an unconditional DT_NEEDED on libhccl.so but imports 0 symbols from it99# WHY stub: torch_npu _C.so has an unconditional DT_NEEDED on libhccl.so but imports 0 symbols from it
85# -- an empty SONAME-only .so satisfies the loader without the real HCCL lib (which ships in ops).100# -- an empty SONAME-only .so satisfies the loader without the real HCCL lib (which ships in ops).
101+# Written through the installer's own `latest` symlink rather than the versioned cann-<v>/ directory it
102+# points at: same inode, but it does not assume the installer keeps naming that directory after
103+# CANN_VERSION, and it is the same path LD_LIBRARY_PATH below resolves.
86RUN printf 'int __hccl_stub;\n' | gcc -shared -fPIC -x c -Wl,-soname,libhccl.so \104RUN printf 'int __hccl_stub;\n' | gcc -shared -fPIC -x c -Wl,-soname,libhccl.so \
87- -o /usr/local/Ascend/cann-${CANN_VERSION}/aarch64-linux/lib64/libhccl.so -105+ -o /usr/local/Ascend/ascend-toolkit/latest/${ARCH}-linux/lib64/libhccl.so -
88 106 
89-# Env = set_env.sh + aarch64-linux/lib64 (WHY the extra path: set_env omits it, but libhccl + its107+# ONE generated env script is the single source of truth for the Ascend environment, and it is where
90-# siblings live there). ENTRYPOINT applies it for `docker run img <cmd>`; bash.bashrc for `docker exec`.108+# ARCH gets baked in -- a JSON ENTRYPOINT cannot expand a build-arg, but a script it sources can carry
109+# one. Three entry paths all funnel through it: ENTRYPOINT (`docker run img <cmd>`), /etc/bash.bashrc
110+# (`docker exec`), /etc/profile.d (login shells, e.g. a runner that does `bash -lc`, whose /etc/profile
111+# would otherwise reset PATH and drop /opt/venv/bin). The sentinel keeps nested shells from stacking
112+# PATH entries. LD_LIBRARY_PATH needs <arch>-linux/lib64 explicitly: set_env.sh omits it, but the
113+# stubbed libhccl and its siblings live there.
91ENV ASCEND_HOME_PATH=/usr/local/Ascend/ascend-toolkit/latest114ENV ASCEND_HOME_PATH=/usr/local/Ascend/ascend-toolkit/latest
92-ENV PATH=/opt/venv/bin:${ASCEND_HOME_PATH}/aarch64-linux/ccec_compiler/bin:${PATH}115+ENV PATH=/opt/venv/bin:${ASCEND_HOME_PATH}/${ARCH}-linux/ccec_compiler/bin:${PATH}
116+# Architecture and CANN version are properties of THIS image, so publish them for anything built FROM
117+# here (docker/eval needs both to pick the right ops .run). A child that re-declared its own ARG could
118+# disagree with the base and silently install the wrong architecture's -- or the wrong release's --
119+# ops on top of this toolkit; inheriting makes that impossible.
120+ENV CANN_ARCH=${ARCH}
121+ENV CANN_VERSION=${CANN_VERSION}
93RUN printf '%s\n' \122RUN printf '%s\n' \
94- 'source /usr/local/Ascend/ascend-toolkit/set_env.sh 2>/dev/null || true' \123+ '# Only set_env.sh is guarded: it appends unboundedly, so re-sourcing it in nested shells' \
95- 'export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/aarch64-linux/lib64:${LD_LIBRARY_PATH}' \124+ '# would grow the environment without limit. The exported vars it defines survive anyway.' \
96- >> /etc/bash.bashrc125+ 'if [ -z "${_CANN_ENV_DONE:-}" ]; then' \
126+ ' export _CANN_ENV_DONE=1' \
127+ ' source /usr/local/Ascend/ascend-toolkit/set_env.sh 2>/dev/null || true' \
128+ 'fi' \
129+ '# PATH / LD_LIBRARY_PATH are re-asserted EVERY time, membership-tested rather than' \
130+ '# guard-tested. Anything downstream may reset PATH after the guard is already exported --' \
131+ "# /etc/profile in a login shell does exactly that -- and a one-shot guard would then skip" \
132+ '# the repair, leaving the shell without /opt/venv/bin (python3 disappears). The case test' \
133+ '# keeps repeated sourcing from growing them.' \
134+ 'case ":${PATH}:" in' \
135+ ' *":/opt/venv/bin:"*) ;;' \
136+ " *) export PATH=/opt/venv/bin:${ASCEND_HOME_PATH}/${ARCH}-linux/ccec_compiler/bin:\${PATH} ;;" \
137+ 'esac' \
138+ "case \":\${LD_LIBRARY_PATH:-}:\" in" \
139+ " *\":${ASCEND_HOME_PATH}/${ARCH}-linux/lib64:\"*) ;;" \
140+ " *) export LD_LIBRARY_PATH=${ASCEND_HOME_PATH}/${ARCH}-linux/lib64:\${LD_LIBRARY_PATH} ;;" \
141+ 'esac' \
142+ > /etc/cann-env.sh \
143+ && chmod 0644 /etc/cann-env.sh \
144+ && ln -sf /etc/cann-env.sh /etc/profile.d/10-cann.sh \
145+ && printf 'source /etc/cann-env.sh\n' >> /etc/bash.bashrc
146+ENV BASH_ENV=/etc/cann-env.sh
97 147 
98-ENTRYPOINT ["/bin/bash", "-c", "source /usr/local/Ascend/ascend-toolkit/set_env.sh 2>/dev/null; export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/aarch64-linux/lib64:${LD_LIBRARY_PATH}; exec \"$@\"", "bash"]148+ENTRYPOINT ["/bin/bash", "-c", "source /etc/cann-env.sh; exec \"$@\"", "bash"]
99WORKDIR /workspace149WORKDIR /workspace
100CMD ["bash"]150CMD ["bash"]
@@ -11,6 +11,7 @@ AscendC/CCE kernel,经 `KNAME<<<grid, nullptr, stream>>>` 直接下发,不走 ac
11| base | AscendHub 完整 CANN(`cann:<ver>-<device>-...`,per-device) | `debian:12-slim`,从 `.run` 自装 toolkit |11| base | AscendHub 完整 CANN(`cann:<ver>-<device>-...`,per-device) | `debian:12-slim`,从 `.run` 自装 toolkit |
12| ops / nnal | 有 | **无(0 ops)** |12| ops / nnal | 有 | **无(0 ops)** |
13| chip | per-`DEVICE` tag | **chip-agnostic**(chip 只进 mounted driver + bisheng `--soc`) |13| chip | per-`DEVICE` tag | **chip-agnostic**(chip 只进 mounted driver + bisheng `--soc`) |
14+| CPU 架构 | per-tag | `ARCH` build-arg(aarch64 / x86_64),原生构建;写入 `ENV CANN_ARCH` 供下游继承 |
14| py / env | ubuntu22.04 + py3.12 | uv 管理的 py3.13 standalone(`uv.lock` 锁定) |15| py / env | ubuntu22.04 + py3.12 | uv 管理的 py3.13 standalone(`uv.lock` 锁定) |
15| 适用 | 全量评测(含 aclnn baseline + perf 开箱) | 直调提交:精度独立可跑;perf 见下 |16| 适用 | 全量评测(含 aclnn baseline + perf 开箱) | 直调提交:精度独立可跑;perf 见下 |
16 17 
@@ -25,10 +26,16 @@ AscendHub per-device 镜像)。
25 26 
26```bash27```bash
27cd docker/base/28cd docker/base/
28-docker build -t cann-toolkit-base:9.0.1-py3.13 .29+docker build --build-arg ARCH=$(uname -m) -t cann-toolkit-base:9.0.1-py3.13 .
29```30```
30 31 
31-镜像 **aarch64-only**(Ascend host = Kunpeng/ARM;x86_64 若需另开专门测过的变体)。python 依赖由32+镜像架构由 **`ARCH` build-arg** 决定(`aarch64` / `x86_64`,默认 `aarch64`),**必须在目标架构的
33+机器上原生构建** —— 没有交叉编译、没有 `--platform`。两种架构都已实测跑通(aarch64/910B2 与
34+x86_64/950PR:装 toolkit、同一份 lock `uv sync --frozen`、stub libhccl、编 `cann_bench_utils`、
35+认卡)。构建结果会把架构写进 `ENV CANN_ARCH`,`docker/eval` **继承**它来挑对应架构的 ops 包 ——
36+所以下游不该、也不需要再声明自己的 `ARCH`。
37+ 
38+python 依赖由
32`pyproject.toml` + `uv.lock` 锁定(hash 校验),`uv sync --frozen` 装入 `/opt/venv`。39`pyproject.toml` + `uv.lock` 锁定(hash 校验),`uv sync --frozen` 装入 `/opt/venv`。
33 40 
34### 镜像源(每个都默认走官方/全球源;受限网络用 `--build-arg` 换在区镜像)41### 镜像源(每个都默认走官方/全球源;受限网络用 `--build-arg` 换在区镜像)
@@ -41,7 +48,32 @@ docker build -t cann-toolkit-base:9.0.1-py3.13 .
41| `UV_PYTHON_INSTALL_MIRROR` | (空 = github releases) | `https://mirror.nju.edu.cn/github-release/astral-sh/python-build-standalone` |48| `UV_PYTHON_INSTALL_MIRROR` | (空 = github releases) | `https://mirror.nju.edu.cn/github-release/astral-sh/python-build-standalone` |
42| `PYPI_MIRROR` | (空 = `files.pythonhosted.org`) | `https://mirrors.huaweicloud.com/repository/pypi` |49| `PYPI_MIRROR` | (空 = `files.pythonhosted.org`) | `https://mirrors.huaweicloud.com/repository/pypi` |
43| `TORCH_MIRROR` | (空 = `download.pytorch.org`) | `https://mirror.nju.edu.cn/pytorch/whl/cpu` |50| `TORCH_MIRROR` | (空 = `download.pytorch.org`) | `https://mirror.nju.edu.cn/pytorch/whl/cpu` |
44-| `CANN_VERSION` | `9.0.1` | toolkit `.run` 版本(从 OBS 拉取) |51+| `CANN_VERSION` | `9.0.1` | toolkit `.run` 版本(从 OBS 拉取);落到 `ENV CANN_VERSION` 供 `docker/eval` 继承 |
52+| `CANN_TOOLKIT_URL` | 空(由 `CANN_VERSION` + `ARCH` 推导) | 只在该版本不按常规命名/不在常规 bucket 时设,见下 |
53+| `ARCH` | `aarch64` | `aarch64` / `x86_64`;决定 toolkit `.run`、`<arch>-linux/` 路径,并落到 `ENV CANN_ARCH` |
54+ 
55+### 换 CANN 版本
56+ 
57+正常只需 `--build-arg CANN_VERSION=<版本>` —— toolkit 和 ops 两个包在
58+`ascend-repo.obs.cn-east-2` 上的命名跨版本一致(9.1.0 的 toolkit/ops × aarch64/x86_64 四个包
59+均已核实存在),`docker/eval` 会继承 `ENV CANN_VERSION` 去推导同版本的 ops:
60+ 
61+```bash
62+# base
63+docker build --build-arg CANN_VERSION=9.1.0 --build-arg ARCH=$(uname -m) \
64+ -t cann-toolkit-base:9.1.0-py3.13 .
65+# eval (build.sh 用 CANN_VERSION 只是为了拼 BASE_IMAGE 的 tag)
66+CANN_VERSION=9.1.0 NPU_ARCH=ascend950 bash ../eval/build.sh
67+```
68+ 
69+`CANN_TOOLKIT_URL` 是给例外准备的:部分版本发在 `ascend-cann-open.obs.cn-north-4` 上,
70+或以合并包 `Ascend-cann_<版本>_linux-<arch>.run` 的形式发布(而非 `Ascend-cann-toolkit_...`)。
71+ 
72+**9.1.0 实测过**(q7 / 910B2 / aarch64):只加 `--build-arg CANN_VERSION=9.1.0`,base + eval
73+(`OPS_MODE=refonly`,顺带验了 9.1.0 的 ops 包)全部建成,`--self-test` 全绿、内置算子仍被挡住,
74+一次真实三阶段评测 4/4 通过、得分 74.10(9.0.1 同一提交同一算子是 73.00)。安装布局也没变
75+(`cann-9.1.0/` + `latest` 符号链接 + `ascend-toolkit/set_env.sh`),`torch_npu 2.10.0.post2`
76+虽然官方对表写的是 CANN 9.0.x,在 9.1.0 上导入、认卡、H2D、编译提交都正常。
45 77 
46`PYPI_MIRROR`/`TORCH_MIRROR` 就地改写 `uv.lock` 里的 canonical wheel URL(**同 hash**,`--frozen` 仍校验),78`PYPI_MIRROR`/`TORCH_MIRROR` 就地改写 `uv.lock` 里的 canonical wheel URL(**同 hash**,`--frozen` 仍校验),
47换源不破坏可复现性。CN 全量示例:79换源不破坏可复现性。CN 全量示例:
@@ -26,13 +26,25 @@ NPU_FLAGS=(
26 --device /dev/davinci_manager26 --device /dev/davinci_manager
27 --device /dev/devmm_svm27 --device /dev/devmm_svm
28 --device /dev/hisi_hdc28 --device /dev/hisi_hdc
29- -v /usr/local/Ascend/driver:/usr/local/Ascend/driver:ro
30- -v /usr/local/dcmi:/usr/local/dcmi:ro
31- -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi:ro
32- -v /etc/ascend_install.info:/etc/ascend_install.info:ro
33 -e LD_LIBRARY_PATH="${DRV}"29 -e LD_LIBRARY_PATH="${DRV}"
34 -e ASCEND_RT_VISIBLE_DEVICES="${ASCEND_RT_VISIBLE_DEVICES:-0}"30 -e ASCEND_RT_VISIBLE_DEVICES="${ASCEND_RT_VISIBLE_DEVICES:-0}"
35)31)
32+# Mount each host path only if it exists AND is the right type: docker CREATES a missing bind-mount
33+# source as a root-owned empty DIRECTORY on the host, which litters a shared box and then keeps being
34+# mounted forever (a5 already carries such an empty /usr/local/bin/npu-smi, which would put a directory
35+# on PATH where an executable belongs). Only the driver tree is universal -- dcmi / npu-smi /
36+# ascend_install.info move with the driver install.
37+maybe_mount() { # $1 = required type (d|f), $2 = host path mounted at the same path in-container
38+ case "$1" in
39+ d) [[ -d "$2" ]] || return 0 ;;
40+ f) [[ -f "$2" ]] || return 0 ;;
41+ esac
42+ NPU_FLAGS+=(-v "$2:$2:ro")
43+}
44+maybe_mount d /usr/local/Ascend/driver
45+maybe_mount d /usr/local/dcmi
46+maybe_mount f /usr/local/bin/npu-smi
47+maybe_mount f /etc/ascend_install.info
36 48 
37case "$MODE" in49case "$MODE" in
38 smoke)50 smoke)
@@ -22,33 +22,63 @@ FROM ${BASE_IMAGE}
22ARG PYPI_MIRROR=22ARG PYPI_MIRROR=
23ARG TORCH_MIRROR=23ARG TORCH_MIRROR=
24ARG UV_PYTHON_INSTALL_MIRROR=24ARG UV_PYTHON_INSTALL_MIRROR=
25-ARG CANN_VERSION=9.0.125+ 
26-ARG ARCH=aarch6426+# This layer is where the per-SoC fork lives; everything architecture-shaped and every version-shaped
27+# thing already happened in the base, and arrives as the inherited ENV CANN_ARCH / CANN_VERSION
28+# (deliberately NOT re-declared here -- local knobs could disagree with the base and quietly install
29+# the wrong architecture's, or the wrong release's, ops on top of its toolkit).
30+# NPU_ARCH picks the SoC cann_bench_utils compiles kernels for; OPS_PKG names that SoC's ops .run
31+# (spelled differently from the compiler flag: ascend910b <-> 910b, ascend950 <-> 950).
32+# docker/eval/build.sh derives NPU_ARCH/OPS_PKG/OPS_MODE for you; set them by hand only for a one-off.
33+ARG NPU_ARCH=ascend910b
34+ARG OPS_PKG=910b
27 35 
28# WHY this exists at all: a submission "cheats" by CALLING a builtin (aclnn<Op> / torch_npu op), which36# WHY this exists at all: a submission "cheats" by CALLING a builtin (aclnn<Op> / torch_npu op), which
29-# dispatches into the 4.2G tbe/kernel BINARY tree. The 151M tbe/impl AscendC SOURCE tree is not a cheat37+# dispatches into the multi-GB tbe/kernel BINARY tree (4.2G on 910B, 5.0G on 950). The ~151M tbe/impl
30-# -- it is a legitimate reference we may want an agent to read.38+# AscendC SOURCE tree is not a cheat -- it is a legitimate reference we may want an agent to read.
31-# none (default): never install ops. The base is already 0-ops, so this costs nothing and blocks39+# none : never install ops -- the base is already ops-less, so this costs nothing. Blocks builtins
32-# every builtin. opp/ still exists (the toolkit ships it, incl. the empty40+# because libopapi (the aclnn entry library, shipped in ops) is absent; the toolkit's own
33-# opp/vendors that .run-form custom-op submissions install into).41+# small ~11M kernel tree is not enough to launch one. opp/ still exists, including the empty
34-# refonly : install ops, then strip ONLY the kernel binaries IN THE SAME LAYER -> keep the42+# opp/vendors that .run-form custom-op submissions install into.
35-# impl source as reference, builtins fail to launch. Use this if a run needs opp43+# refonly : install ops, then strip ONLY the kernel binaries IN THE SAME LAYER -> keep the impl source
36-# machinery that 0-ops lacks (see README's perf note).44+# as reference; libopapi is present but every builtin fails to launch.
37-# full : install ops and keep everything. Cheatable; for re-collecting aclnn baselines.45+# full : install ops and keep everything. Cheatable; for re-collecting aclnn baselines.
46+#
47+# WHICH ONE: `none` is right on 910B and WRONG on 950. Measured on a 950PR: with no ops at all even a
48+# bare `torch.arange(8).npu()` dies -- the H2D copy routes through aclnnInplaceCopy there (ERR01007
49+# "OPS feature not supported"), where 910B does a plain aclrtMemcpy and needs nothing. `refonly` fixes
50+# that on 950 and still blocks builtins (matmul dies 561103 "Parse dynamic kernel config fail"), so 950
51+# keeps the full anti-cheat posture -- it just cannot use `none`. The guard below refuses that combo
52+# rather than shipping an image that dies on its first tensor.
38ARG OPS_MODE=none53ARG OPS_MODE=none
39-ARG CANN_OPS_URL=https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/CANN/CANN%20${CANN_VERSION}/Ascend-cann-910b-ops_${CANN_VERSION}_linux-${ARCH}.run54+# Empty = derive from the inherited CANN_VERSION + CANN_ARCH below. Set it only to pin an exact .run
55+# (the ops naming has held across releases -- 9.1.0 is published for 910b/950 x aarch64/x86_64 under
56+# exactly this scheme -- so this is a pin, not a portability crutch).
57+ARG CANN_OPS_URL=
40ARG OPP_PATH=/usr/local/Ascend/ascend-toolkit/latest/opp58ARG OPP_PATH=/usr/local/Ascend/ascend-toolkit/latest/opp
41 59 
42# The ops tree is GBs and never changes -- keep it as the first (bottom-most) layer so the python /60# The ops tree is GBs and never changes -- keep it as the first (bottom-most) layer so the python /
43# harness layers below rebuild independently of it.61# harness layers below rebuild independently of it.
44RUN set -e; \62RUN set -e; \
45 case "${OPS_MODE}" in none|refonly|full) ;; *) echo "invalid OPS_MODE=${OPS_MODE}" >&2; exit 1;; esac; \63 case "${OPS_MODE}" in none|refonly|full) ;; *) echo "invalid OPS_MODE=${OPS_MODE}" >&2; exit 1;; esac; \
64+ : "${CANN_ARCH:?base image did not set CANN_ARCH -- rebuild docker/base}"; \
65+ : "${CANN_VERSION:?base image did not set CANN_VERSION -- rebuild docker/base}"; \
66+ if [ "$(uname -m)" != "${CANN_ARCH}" ]; then \
67+ echo "base image is ${CANN_ARCH} but this is a $(uname -m) host -- native builds only." >&2; \
68+ exit 1; \
69+ fi; \
70+ if [ "${NPU_ARCH}" = ascend950 ] && [ "${OPS_MODE}" = none ]; then \
71+ echo "OPS_MODE=none is not usable on ${NPU_ARCH}: the H2D copy needs aclnnInplaceCopy (ERR01007)." >&2; \
72+ echo "Use OPS_MODE=refonly -- it keeps the anti-cheat posture on 950. See the note above." >&2; \
73+ exit 1; \
74+ fi; \
75+ OPS_URL="${CANN_OPS_URL:-https://ascend-repo.obs.cn-east-2.myhuaweicloud.com/CANN/CANN%20${CANN_VERSION}/Ascend-cann-${OPS_PKG}-ops_${CANN_VERSION}_linux-${CANN_ARCH}.run}"; \
46 if [ "${OPS_MODE}" != none ]; then \76 if [ "${OPS_MODE}" != none ]; then \
47- wget --quiet --header="Referer: https://www.hiascend.com/" "${CANN_OPS_URL}" -O /tmp/ops.run \77+ wget --quiet --header="Referer: https://www.hiascend.com/" "${OPS_URL}" -O /tmp/ops.run \
48 && chmod +x /tmp/ops.run && /tmp/ops.run --quiet --install --install-for-all && rm -f /tmp/ops.run; \78 && chmod +x /tmp/ops.run && /tmp/ops.run --quiet --install --install-for-all && rm -f /tmp/ops.run; \
49 fi; \79 fi; \
50 if [ "${OPS_MODE}" = refonly ]; then rm -rf "${OPP_PATH}/built-in/op_impl/ai_core/tbe/kernel"; fi; \80 if [ "${OPS_MODE}" = refonly ]; then rm -rf "${OPP_PATH}/built-in/op_impl/ai_core/tbe/kernel"; fi; \
51- echo "OPS_MODE=${OPS_MODE}"81+ echo "OPS_MODE=${OPS_MODE} NPU_ARCH=${NPU_ARCH} CANN_ARCH=${CANN_ARCH} CANN_VERSION=${CANN_VERSION}"
52ENV OPS_MODE=${OPS_MODE}82ENV OPS_MODE=${OPS_MODE}
53 83 
54# kernel_eval imports more than the direct-launch base needed (pandas / ruamel.yaml / protobuf /84# kernel_eval imports more than the direct-launch base needed (pandas / ruamel.yaml / protobuf /
@@ -56,9 +86,17 @@ ENV OPS_MODE=${OPS_MODE}
56# `uv sync` only ADDS -- it reconciles the same /opt/venv without uninstalling the base's torch.86# `uv sync` only ADDS -- it reconciles the same /opt/venv without uninstalling the base's torch.
57# Mirror handling is the base's trick: rewrite the lock's canonical URLs in place, same hashes, so87# Mirror handling is the base's trick: rewrite the lock's canonical URLs in place, same hashes, so
58# --frozen still verifies and mirroring costs no reproducibility.88# --frozen still verifies and mirroring costs no reproducibility.
89+#
90+# UV_INDEX_URL is needed ON TOP of that rewrite, and only here. `en-dtypes` is the one entry in this
91+# lock published as an sdist with no wheel, so uv must BUILD it -- and building means resolving its
92+# build-system.requires (setuptools, numpy) from an INDEX, which the URL rewrite does not touch. Left
93+# unset that resolution goes to pypi.org and times out in-region ("Failed to fetch
94+# https://pypi.org/simple/setuptools/ ... operation timed out"), failing the whole layer after ~4min.
95+# docker/base's lock has no sdist-only entry, which is why it never needed this.
59COPY docker/eval/pyproject.toml docker/eval/uv.lock /opt/eval-env/96COPY docker/eval/pyproject.toml docker/eval/uv.lock /opt/eval-env/
60RUN L=/opt/eval-env/uv.lock; \97RUN L=/opt/eval-env/uv.lock; \
61- if [ -n "${PYPI_MIRROR}" ]; then sed -i "s|https://files.pythonhosted.org/packages/|${PYPI_MIRROR%/}/packages/|g" "$L"; fi; \98+ if [ -n "${PYPI_MIRROR}" ]; then sed -i "s|https://files.pythonhosted.org/packages/|${PYPI_MIRROR%/}/packages/|g" "$L"; \
99+ export UV_INDEX_URL="${PYPI_MIRROR%/}/simple"; fi; \
62 if [ -n "${TORCH_MIRROR}" ]; then sed -i "s|https://download-r2.pytorch.org/whl/cpu/|${TORCH_MIRROR%/}/|g" "$L"; fi; \100 if [ -n "${TORCH_MIRROR}" ]; then sed -i "s|https://download-r2.pytorch.org/whl/cpu/|${TORCH_MIRROR%/}/|g" "$L"; fi; \
63 UV_PYTHON_INSTALL_MIRROR="${UV_PYTHON_INSTALL_MIRROR}" uv sync --frozen --project /opt/eval-env101 UV_PYTHON_INSTALL_MIRROR="${UV_PYTHON_INSTALL_MIRROR}" uv sync --frozen --project /opt/eval-env
64 102 
@@ -75,8 +113,9 @@ ENV CANN_BENCH_DIR=/opt/cann-bench
75# for the frequency ramp, so run_evaluation.sh's ensure_cann_bench_utils() would try to compile it at113# for the frequency ramp, so run_evaluation.sh's ensure_cann_bench_utils() would try to compile it at
76# RUN time. Compile and install it here instead -- it needs only bisheng, no NPU -- so the container114# RUN time. Compile and install it here instead -- it needs only bisheng, no NPU -- so the container
77# starts straight into evaluating and the wheel is part of the frozen artifact.115# starts straight into evaluating and the wheel is part of the frozen artifact.
78-# It embeds SoC-specific kernels, hence NPU_ARCH: this image is per-SoC even though the base is not.116+# It embeds SoC-specific kernels (NPU_ARCH, declared with the other knobs at the top) -- which is what
79-ARG NPU_ARCH=ascend910b117+# makes this image per-SoC even though the base is not. Verified to build and run for both ascend910b
118+# and ascend950.
80RUN . /usr/local/Ascend/ascend-toolkit/set_env.sh 2>/dev/null || true; \119RUN . /usr/local/Ascend/ascend-toolkit/set_env.sh 2>/dev/null || true; \
81 cd /opt/cann-bench/src/cann_bench_utils \120 cd /opt/cann-bench/src/cann_bench_utils \
82 && PYTHON=/opt/venv/bin/python3 bash build.sh --clean --soc=${NPU_ARCH} \121 && PYTHON=/opt/venv/bin/python3 bash build.sh --clean --soc=${NPU_ARCH} \
@@ -8,35 +8,64 @@ harness 冻结在镜像里(`src/kernel_eval` + `tasks/` + `cann_bench_utils` 全
8| `/submission` | AI 生成的算子源码目录 | 只读即可(entrypoint 会先复制) |8| `/submission` | AI 生成的算子源码目录 | 只读即可(entrypoint 会先复制) |
9| `/reports` | 评测报告 + `prof_data/` + `build/`(编译日志、wheel) | 读写 |9| `/reports` | 评测报告 + `prof_data/` + `build/`(编译日志、wheel) | 读写 |
10 10 
11-**镜像 tag 就是 benchmark 版本**:`cann-bench-eval:<VERSION>-<NPU_ARCH>-ops<OPS_MODE>`,11+**镜像 tag 就是 benchmark 版本 + 目标**:
12-例如 `cann-bench-eval:1.0.0-ascend910b-opsnone`。12+`cann-bench-eval:<VERSION>-<NPU_ARCH>-<ARCH>-ops<OPS_MODE>`,
13+例如 `cann-bench-eval:1.0.0-ascend910b-aarch64-opsnone`。
14+ 
15+## 分工:哪一层管什么
16+ 
17+`docker/base` 就是 common —— 它到 toolkit 为止,**不装 ops**,唯一的变量是 CPU 架构。
18+**SoC 分叉正好从 `ops.run` 开始**,所以整条分叉都在本层,而且只有三个值:
19+ 
20+| 轴 | 归属 | 怎么定 |
21+|---|---|---|
22+| CPU 架构(aarch64 / x86_64) | **base** | base 的 `ARCH` build-arg;建完写进 `ENV CANN_ARCH`,本层**继承**它,不再声明自己的 —— 否则两边可以不一致,悄悄往 arm 镜像装 x86 的 ops |
23+| SoC(910b / 910_93 / 950) | **eval** | `NPU_ARCH`(给 `cann_bench_utils` 编 kernel)+ `OPS_PKG`(ops 包名)+ `OPS_MODE` |
24+ 
25+因此单 Dockerfile 就够:a2 与 a5 镜像的差别只有 4 个 ARG 值,**零结构差异**。
13 26 
14## Build27## Build
15 28 
16-底座是 [`docker/base`](../base/) 的 `cann-toolkit-base`,先有它:29+底座是 [`docker/base`](../base/) 的 `cann-toolkit-base`,先有它(**同一台机器、同一架构**,
30+这是原生构建,不是交叉编译):
17 31 
18```bash32```bash
19-cd docker/base && docker build -t cann-toolkit-base:9.0.1-py3.13 .33+cd docker/base && docker build --build-arg ARCH=$(uname -m) -t cann-toolkit-base:9.0.1-py3.13 .
20```34```
21 35 
22然后(**build context 必须是仓库根**,`build.sh` 已经处理好):36然后(**build context 必须是仓库根**,`build.sh` 已经处理好):
23 37 
24```bash38```bash
25-bash docker/eval/build.sh # 默认: OPS_MODE=none, ascend910b39+bash docker/eval/build.sh # 本机架构 + ascend910b + OPS_MODE=none
26-OPS_MODE=refonly bash docker/eval/build.sh # 见下"ops 模式"40+NPU_ARCH=ascend950 bash docker/eval/build.sh # 950PR —— OPS_MODE 自动取 refonly,见下
27NPU_ARCH=ascend910_93 bash docker/eval/build.sh # A341NPU_ARCH=ascend910_93 bash docker/eval/build.sh # A3
28MIRROR=cn bash docker/eval/build.sh # 受限网络: 一把切到在区镜像源42MIRROR=cn bash docker/eval/build.sh # 受限网络: 一把切到在区镜像源
29TRITON_ASCEND_VERSION=3.2.1 bash docker/eval/build.sh43TRITON_ASCEND_VERSION=3.2.1 bash docker/eval/build.sh
30```44```
31 45 
46+`build.sh` 由 `NPU_ARCH` + `uname -m` 推导其余一切,正常情况下你只需要说芯片:
47+ 
48+| `NPU_ARCH` | → `OPS_PKG` | → 默认 `OPS_MODE` |
49+|---|---|---|
50+| `ascend910b` | `910b` | `none` |
51+| `ascend910_93` | `910_93` | `none` |
52+| `ascend950` | `950` | **`refonly`**(`none` 在 950 上不可用,见下) |
53+ 
32| build-arg | 默认 | 说明 |54| build-arg | 默认 | 说明 |
33|---|---|---|55|---|---|---|
34-| `BASE_IMAGE` | `cann-toolkit-base:9.0.1-py3.13` | 底座 |56+| `BASE_IMAGE` | `cann-toolkit-base:<CANN_VERSION>-py3.13` | 底座;其 `CANN_ARCH` 必须与本机架构一致,`build.sh` 会校验 |
35-| `OPS_MODE` | `none` | `none` / `refonly` / `full`,见下 |
36| `NPU_ARCH` | `ascend910b` | `cann_bench_utils` 的 kernel 是 SoC 相关的,故本镜像 per-SoC |57| `NPU_ARCH` | `ascend910b` | `cann_bench_utils` 的 kernel 是 SoC 相关的,故本镜像 per-SoC |
58+| `OPS_PKG` | `910b` | ops `.run` 的 SoC 拼写(与编译器 flag 不同:`ascend950` ↔ `950`) |
59+| `OPS_MODE` | 见上表 | `none` / `refonly` / `full`,见下 |
60+| `CANN_OPS_URL` | 空(由 `CANN_VERSION` + `CANN_ARCH` + `OPS_PKG` 推导) | 只在需要钉死某个 `.run` 时设 |
37| `TRITON_ASCEND_VERSION` | 空 | 非空则装 Triton-Ascend(体积大),语义同 `docker/dev` |61| `TRITON_ASCEND_VERSION` | 空 | 非空则装 Triton-Ascend(体积大),语义同 `docker/dev` |
38| `PYPI_MIRROR` / `TORCH_MIRROR` / `UV_PYTHON_INSTALL_MIRROR` | 空(官方源) | 同 `docker/base`;`MIRROR=cn` 是这三个的快捷方式 |62| `PYPI_MIRROR` / `TORCH_MIRROR` / `UV_PYTHON_INSTALL_MIRROR` | 空(官方源) | 同 `docker/base`;`MIRROR=cn` 是这三个的快捷方式 |
39 63 
64+**没有 `ARCH` / `CANN_VERSION` build-arg** —— 这两样都由 base 经 `ENV` 单向下传:架构决定装哪个
65+架构的 ops,版本决定装哪个 release 的 ops,本层各自重新声明就可能和底座不一致。base 缺
66+`CANN_ARCH` / `CANN_VERSION`(即早于跨架构改动)、或 base 架构与本机不符,build 期三道断言都会
67+直接失败。换 CANN 版本只需重建 base,见 `docker/base/README.md`。
68+ 
40python 依赖由 `docker/eval/{pyproject.toml,uv.lock}` 锁定,是 `docker/base` 依赖集的**严格超集**69python 依赖由 `docker/eval/{pyproject.toml,uv.lock}` 锁定,是 `docker/base` 依赖集的**严格超集**
41(`uv sync` 会把环境对齐到 lock —— 非超集会把底座已装的 torch 卸掉)。改任一边都要同步另一边。70(`uv sync` 会把环境对齐到 lock —— 非超集会把底座已装的 torch 卸掉)。改任一边都要同步另一边。
42 71 
@@ -98,23 +127,38 @@ docker run ... -v "$PWD:/opt/cann-bench" cann-bench-eval:... /submission --opera
98 127 
99## ops 模式(反作弊形态)128## ops 模式(反作弊形态)
100 129 
101-提交"作弊"的方式是**调用内置算子**(`aclnn<Op>` / `torch_npu` op),它们下发到 4.2G 的130+提交"作弊"的方式是**调用内置算子**(`aclnn<Op>` / `torch_npu` op),它们下发到
102-`opp/built-in/op_impl/ai_core/tbe/kernel` 二进制树。旁边 151M 的 `tbe/impl` 是 AscendC **源码**,131+`opp/built-in/op_impl/ai_core/tbe/kernel` 这棵多 GB 的二进制树(910B 4.2G,950 5.0G)。旁边约
103-那不是作弊,是合法参考。132+151M 的 `tbe/impl` 是 AscendC **源码**,那不是作弊,是合法参考。
104 133 
105| `OPS_MODE` | 做什么 | 后果 |134| `OPS_MODE` | 做什么 | 后果 |
106|---|---|---|135|---|---|---|
107-| `none`(默认) | 不装 ops(底座本来就是 0-ops) | 内置算子根本不存在,无从蹭起;镜像最小。`opp/` 目录仍在(toolkit 自带),`.run` 形态的自定义算子提交照常装进 `opp/vendors` |136+| `none` | 不装 ops(底座本来就没有) | 内置算子起不来 —— 挡住它的是 **`libopapi` 缺席**(aclnn 的入口库在 ops 包里),而不是"kernel 树为空":toolkit 自带一棵 ~11M 的 `tbe/kernel`,但那不足以下发内置算子。镜像最小。`opp/` 仍在,`.run` 形态的自定义算子提交照常装进 `opp/vendors` |
108-| `refonly` | 装 ops,**同层**删掉 `tbe/kernel` 二进制 | 保留 `tbe/impl` AscendC 源码作参考,内置算子下发失败;比 `none` 多出 opp 的全套机制 |137+| `refonly` | 装 ops,**同层**删掉 `tbe/kernel` 二进制 | 保留 `tbe/impl` AscendC 源码作参考;`libopapi` 在,但内置算子下发失败。比 `none` 多出 opp 的全套机制 |
109| `full` | 装 ops 不删 | 可被蹭内置算子;用于重采 aclnn baseline |138| `full` | 装 ops 不删 | 可被蹭内置算子;用于重采 aclnn baseline |
110 139 
111-自检的 `[6]` 项会直说当前镜像里内置算子能不能下发。140+自检的 `[6]` 项会直说当前镜像里内置算子能不能下发,报错码还能区分是哪种形态:
141+`none` → `500001 LazyInitAclops`;`refonly` → `561103 Parse dynamic kernel config fail`。
112 142 
113-**默认 `none` 够跑全量三阶段(编译/精度/性能)** —— 910B2 上实测:`direct_launch_example` 的 Sqrt143+### 选哪个:910B 用 `none`,950 **必须** `refonly`
144+ 
145+| | 910B2 / aarch64 / `none` | 950PR / x86_64 / `refonly` |
146+|---|---|---|
147+| 裸 `.npu()` H2D 拷贝 | PASS | PASS |
148+| `cann_bench_warmup`(10240²) | ok 4.5 ms | ok 6.4 ms |
149+| `cann_bench_cache_clean`(96×1024²) | ok 0.7 ms | ok 1.1 ms |
150+| 内置 `matmul` | BLOCKED(500001) | BLOCKED(561103) |
151+ 
152+**`none` 在 950 上不可用**:实测那里连一次 `torch.arange(8).npu()` 都会死 —— H2D 拷贝在 950 上
153+走 `aclnnInplaceCopy`(`ERR01007 OPS feature not supported`),而 910B 走的是普通 `aclrtMemcpy`,
154+什么都不需要。`refonly` 修好这条且**仍然挡住内置算子**,所以 950 的反作弊形态是完整的,只是不能用
155+`none`。Dockerfile 里有 build 期断言直接拒绝 `ascend950 + none` 这个组合,不会让一个"第一个张量
156+就崩"的镜像出厂。
157+ 
158+**910B 上默认 `none` 够跑全量三阶段(编译/精度/性能)** —— 实测 `direct_launch_example` 的 Sqrt
1144/4 精度通过,profiler 正常产出 `prof_data/` 与 device kernel 耗时(`sqrt_kernel` 6.9us),综合得分1594/4 精度通过,profiler 正常产出 `prof_data/` 与 device kernel 耗时(`sqrt_kernel` 6.9us),综合得分
11573.00。`LazyInitAclops` 在 0-ops 下确实会失败(自检 `[6]` 就是它),但性能采集不经过这条路 ——16073.00。`LazyInitAclops` 在 0-ops 下确实会失败(自检 `[6]` 就是它),但性能采集不经过这条路 ——
116-升频/清 cache 由镜像里烘好的 `cann_bench_utils` 直调 kernel 提供。所以 `refonly` 只在提交本身161+升频/清 cache 由镜像里烘好的 `cann_bench_utils` 直调 kernel 提供。
117-需要 opp 全套机制时才用得上。
118 162 
119## 已知取舍163## 已知取舍
120 164 
@@ -123,3 +167,7 @@ docker run ... -v "$PWD:/opt/cann-bench" cann-bench-eval:... /submission --opera
123- `tasks/` 烘进镜像,所以镜像 tag 即 benchmark 版本;开发期用上面的 `-v` 覆盖回工作树。167- `tasks/` 烘进镜像,所以镜像 tag 即 benchmark 版本;开发期用上面的 `-v` 覆盖回工作树。
124- `cann_bench_utils` 在 build 期就编好装好(只要 bisheng,不需要 NPU),容器起来直接开跑,168- `cann_bench_utils` 在 build 期就编好装好(只要 bisheng,不需要 NPU),容器起来直接开跑,
125 `ensure_cann_bench_utils()` 短路返回。它含 SoC 相关 kernel,所以本镜像 per-SoC。169 `ensure_cann_bench_utils()` 短路返回。它含 SoC 相关 kernel,所以本镜像 per-SoC。
170+- **原生构建,没有交叉编译**:base 和 eval 必须在目标架构的机器上建。`uname -m` 与 base 的
171+ `CANN_ARCH` 不符时 build 期直接失败,而不是产出一个跑不起来的镜像。
172+- 在区网络下 `MIRROR=cn` 基本是必需项,不是可选项:不换源时 `uv sync` 会卡在 PyPI 上几十分钟
173+ (纯网络阻塞,看着像挂死);a5 上 apt 和 docker.io 同样需要换源。
@@ -2,9 +2,10 @@
2# Build the cann-bench-eval image. Run from anywhere; the build context is forced to the repo root2# Build the cann-bench-eval image. Run from anywhere; the build context is forced to the repo root
3# because the Dockerfile COPYs src/ tasks/ scripts/ (see .dockerignore for what is kept out).3# because the Dockerfile COPYs src/ tasks/ scripts/ (see .dockerignore for what is kept out).
4#4#
5-# bash docker/eval/build.sh # OPS_MODE=none (default), 910b5+# bash docker/eval/build.sh # this host's arch, 910b, OPS_MODE=none
6-# OPS_MODE=refonly bash docker/eval/build.sh # ops installed, kernel binaries stripped6+# NPU_ARCH=ascend950 bash docker/eval/build.sh # 950PR -- OPS_MODE defaults to refonly (see below)
7# NPU_ARCH=ascend910_93 bash docker/eval/build.sh # A37# NPU_ARCH=ascend910_93 bash docker/eval/build.sh # A3
8+# OPS_MODE=full bash docker/eval/build.sh # cheatable; for re-collecting aclnn baselines
8# TRITON_ASCEND_VERSION=3.2.1 bash docker/eval/build.sh9# TRITON_ASCEND_VERSION=3.2.1 bash docker/eval/build.sh
9#10#
10# CN mirrors (see README): MIRROR=cn bash docker/eval/build.sh11# CN mirrors (see README): MIRROR=cn bash docker/eval/build.sh
@@ -13,21 +14,43 @@ set -euo pipefail
13REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"14REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
14cd "$REPO_ROOT"15cd "$REPO_ROOT"
15 16 
16-BASE_IMAGE="${BASE_IMAGE:-cann-toolkit-base:9.0.1-py3.13}"17+# ARCH is the BUILD HOST's -- this is a native build, not a cross-compile. Overriding it to something
17-OPS_MODE="${OPS_MODE:-none}"18+# other than `uname -m` produces an image that cannot run here. Normalise the aliases: CANN's .run
19+# names and the toolkit's <arch>-linux dirs use aarch64/x86_64, while uname says arm64 on macOS and
20+# amd64 is the common docker spelling -- an unnormalised value silently builds a 404 download URL.
21+ARCH="${ARCH:-$(uname -m)}"
22+case "${ARCH}" in
23+ arm64|aarch64) ARCH=aarch64 ;;
24+ amd64|x86_64) ARCH=x86_64 ;;
25+ *) echo "unsupported ARCH=${ARCH} (expected aarch64 | x86_64)" >&2; exit 1 ;;
26+esac
18NPU_ARCH="${NPU_ARCH:-ascend910b}"27NPU_ARCH="${NPU_ARCH:-ascend910b}"
19CANN_VERSION="${CANN_VERSION:-9.0.1}"28CANN_VERSION="${CANN_VERSION:-9.0.1}"
29+BASE_IMAGE="${BASE_IMAGE:-cann-toolkit-base:${CANN_VERSION}-py3.13}"
20TRITON_ASCEND_VERSION="${TRITON_ASCEND_VERSION:-}"30TRITON_ASCEND_VERSION="${TRITON_ASCEND_VERSION:-}"
21 31 
22-VERSION="$(cat VERSION)"32+# The ops .run spells the SoC differently from the compiler flag (ascend910b -> 910b, ascend950 -> 950).
23-# Tag carries everything that changes what a score means: benchmark version, ops posture, SoC.33+case "${NPU_ARCH}" in
24-IMAGE="${IMAGE:-cann-bench-eval:${VERSION}-${NPU_ARCH}-ops${OPS_MODE}}"34+ ascend910b) OPS_PKG="${OPS_PKG:-910b}" ; DEFAULT_OPS_MODE=none ;;
35+ ascend910_93) OPS_PKG="${OPS_PKG:-910_93}" ; DEFAULT_OPS_MODE=none ;;
36+ # 950 cannot run OPS_MODE=none -- even a bare .npu() copy needs aclnnInplaceCopy there (ERR01007).
37+ # refonly still blocks builtins, so the anti-cheat posture is preserved. See docker/eval/README.md.
38+ ascend950) OPS_PKG="${OPS_PKG:-950}" ; DEFAULT_OPS_MODE=refonly ;;
39+ *) echo "unknown NPU_ARCH=${NPU_ARCH} (expected ascend910b | ascend910_93 | ascend950)" >&2; exit 1 ;;
40+esac
41+OPS_MODE="${OPS_MODE:-${DEFAULT_OPS_MODE}}"
25 42 
43+VERSION="$(cat VERSION)"
44+# Tag carries everything that changes what a score means: benchmark version, SoC, CPU arch, ops posture.
45+IMAGE="${IMAGE:-cann-bench-eval:${VERSION}-${NPU_ARCH}-${ARCH}-ops${OPS_MODE}}"
46+ 
47+# No ARCH / CANN_VERSION build-arg: the eval layer inherits CANN_ARCH and CANN_VERSION from the base
48+# image. Both are used here only for the tag, the base-image name, and the consistency check below.
26ARGS=(49ARGS=(
27 --build-arg "BASE_IMAGE=${BASE_IMAGE}"50 --build-arg "BASE_IMAGE=${BASE_IMAGE}"
28 --build-arg "OPS_MODE=${OPS_MODE}"51 --build-arg "OPS_MODE=${OPS_MODE}"
52+ --build-arg "OPS_PKG=${OPS_PKG}"
29 --build-arg "NPU_ARCH=${NPU_ARCH}"53 --build-arg "NPU_ARCH=${NPU_ARCH}"
30- --build-arg "CANN_VERSION=${CANN_VERSION}"
31 --build-arg "TRITON_ASCEND_VERSION=${TRITON_ASCEND_VERSION}"54 --build-arg "TRITON_ASCEND_VERSION=${TRITON_ASCEND_VERSION}"
32)55)
33 56 
@@ -41,11 +64,30 @@ if [[ "${MIRROR:-}" == "cn" ]]; then
41fi64fi
42 65 
43if ! docker image inspect "${BASE_IMAGE}" >/dev/null 2>&1; then66if ! docker image inspect "${BASE_IMAGE}" >/dev/null 2>&1; then
44- echo "==> base image ${BASE_IMAGE} not found; build it first: cd docker/base && docker build -t ${BASE_IMAGE} ." >&267+ echo "==> base image ${BASE_IMAGE} not found. Build it first (same ARCH, same host):" >&2
68+ echo " cd docker/base && docker build --build-arg ARCH=${ARCH} -t ${BASE_IMAGE} ." >&2
45 exit 169 exit 1
46fi70fi
47 71 
48-echo "==> building ${IMAGE} (base=${BASE_IMAGE} ops=${OPS_MODE} soc=${NPU_ARCH})"72+# The base owns the architecture and the CANN release; fail loudly here rather than let the tag claim
73+# one thing while the image is another. (The Dockerfile re-checks the arch against `uname -m` inside
74+# the build, and refuses a base that publishes neither variable.)
75+BASE_ENV="$(docker image inspect "${BASE_IMAGE}" --format '{{range .Config.Env}}{{println .}}{{end}}')"
76+BASE_ARCH="$(sed -n 's/^CANN_ARCH=//p' <<<"${BASE_ENV}")"
77+BASE_CANN="$(sed -n 's/^CANN_VERSION=//p' <<<"${BASE_ENV}")"
78+if [[ -z "${BASE_ARCH}" || -z "${BASE_CANN}" ]]; then
79+ MISSING=""
80+ [[ -z "${BASE_ARCH}" ]] && MISSING="CANN_ARCH"
81+ [[ -z "${BASE_CANN}" ]] && MISSING="${MISSING:+${MISSING} }CANN_VERSION"
82+ echo "==> ${BASE_IMAGE} publishes no ${MISSING} -- it predates the cross-arch change; rebuild docker/base." >&2
83+ exit 1
84+fi
85+if [[ "${BASE_ARCH}" != "${ARCH}" ]]; then
86+ echo "==> ${BASE_IMAGE} is ${BASE_ARCH} but this host is ${ARCH}; rebuild the base here." >&2
87+ exit 1
88+fi
89+ 
90+echo "==> building ${IMAGE} (base=${BASE_IMAGE} cann=${BASE_CANN} arch=${ARCH} soc=${NPU_ARCH} ops=${OPS_MODE}/${OPS_PKG})"
49set -x91set -x
50docker build --network=host -f docker/eval/Dockerfile -t "${IMAGE}" "${ARGS[@]}" "$@" .92docker build --network=host -f docker/eval/Dockerfile -t "${IMAGE}" "${ARGS[@]}" "$@" .
51set +x93set +x
@@ -8,14 +8,14 @@ SUBMISSION_DIR="${SUBMISSION_DIR:-/submission}"
8REPORTS_DIR="${REPORTS_DIR:-/reports}"8REPORTS_DIR="${REPORTS_DIR:-/reports}"
9WORK_SRC=/work/src9WORK_SRC=/work/src
10 10 
11-# The base image's ENTRYPOINT did this; we replaced it, so redo it here. ENV covers PATH, but11+# The base image's ENTRYPOINT sourced /etc/cann-env.sh; this image replaces that ENTRYPOINT, so redo
12-# set_env.sh has side effects (ASCEND_OPP_PATH, ASCEND_AICPU_PATH) the eval and the .run-form12+# it. That script is the base's single source of truth for the Ascend environment: set_env.sh (whose
13-# custom-op install depend on, and lib64 is missing from set_env's own LD_LIBRARY_PATH.13+# ASCEND_OPP_PATH / ASCEND_AICPU_PATH side effects the eval and the .run-form custom-op install depend
14-source /usr/local/Ascend/ascend-toolkit/set_env.sh 2>/dev/null || true14+# on), the venv on PATH, and the <arch>-linux/lib64 that set_env.sh itself omits -- with the image's
15-export LD_LIBRARY_PATH=/usr/local/Ascend/ascend-toolkit/latest/aarch64-linux/lib64:${LD_LIBRARY_PATH:-}15+# own architecture already baked in. Re-deriving any of it here is what left an aarch64 lib64 path
16-# set_env.sh rewrites PATH; re-assert the venv in front of it. There is no system python in the16+# hardcoded in an image that also ships for x86_64. Sourcing is idempotent, so repeating what BASH_ENV
17-# debian-slim base, so losing /opt/venv/bin means run_evaluation.sh's `command -v python` fails.17+# already did costs nothing.
18-export PATH=/opt/venv/bin:${PATH}18+source /etc/cann-env.sh
19 19 
20usage() {20usage() {
21 cat <<EOF21 cat <<EOF
@@ -16,9 +16,12 @@ set -euo pipefail
16REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"16REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
17 17 
18VERSION="$(cat "${REPO_ROOT}/VERSION")"18VERSION="$(cat "${REPO_ROOT}/VERSION")"
19+ARCH="${ARCH:-$(uname -m)}"
20+case "${ARCH}" in arm64|aarch64) ARCH=aarch64 ;; amd64|x86_64) ARCH=x86_64 ;; esac
19NPU_ARCH="${NPU_ARCH:-ascend910b}"21NPU_ARCH="${NPU_ARCH:-ascend910b}"
20-OPS_MODE="${OPS_MODE:-none}"22+# Mirrors build.sh's per-SoC default (950 cannot run OPS_MODE=none) so the derived tag matches.
21-IMAGE="${IMAGE:-cann-bench-eval:${VERSION}-${NPU_ARCH}-ops${OPS_MODE}}"23+case "${NPU_ARCH}" in ascend950) OPS_MODE="${OPS_MODE:-refonly}" ;; *) OPS_MODE="${OPS_MODE:-none}" ;; esac
24+IMAGE="${IMAGE:-cann-bench-eval:${VERSION}-${NPU_ARCH}-${ARCH}-ops${OPS_MODE}}"
22REPORTS="${REPORTS:-${PWD}/reports}"25REPORTS="${REPORTS:-${PWD}/reports}"
23 26 
24DRV=/usr/local/Ascend/driver/lib64/driver:/usr/local/Ascend/driver/lib6427DRV=/usr/local/Ascend/driver/lib64/driver:/usr/local/Ascend/driver/lib64
@@ -29,12 +32,25 @@ NPU_FLAGS=(
29 --device /dev/davinci_manager32 --device /dev/davinci_manager
30 --device /dev/devmm_svm33 --device /dev/devmm_svm
31 --device /dev/hisi_hdc34 --device /dev/hisi_hdc
32- -v /usr/local/Ascend/driver:/usr/local/Ascend/driver:ro
33- -v /usr/local/dcmi:/usr/local/dcmi:ro
34- -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi:ro
35- -v /etc/ascend_install.info:/etc/ascend_install.info:ro
36 -e LD_LIBRARY_PATH="${DRV}"35 -e LD_LIBRARY_PATH="${DRV}"
37)36)
37+# Only the driver tree is universal. dcmi / npu-smi / ascend_install.info sit wherever the host's
38+# driver install put them, and docker CREATES a missing bind-mount source as a root-owned empty
39+# DIRECTORY on the host -- littering a shared box and shadowing the in-container path. Mount each only
40+# if it exists AND is the right type: a5 already carries an empty /usr/local/bin/npu-smi directory left
41+# by some earlier unconditional mount, and passing that through would put a directory on PATH where an
42+# executable belongs. npu-smi matters because some submissions' build.sh shells out to it for the SoC.
43+maybe_mount() { # $1 = required type (d|f), $2 = host path mounted at the same path in-container
44+ case "$1" in
45+ d) [[ -d "$2" ]] || return 0 ;;
46+ f) [[ -f "$2" ]] || return 0 ;;
47+ esac
48+ NPU_FLAGS+=(-v "$2:$2:ro")
49+}
50+maybe_mount d /usr/local/Ascend/driver
51+maybe_mount d /usr/local/dcmi
52+maybe_mount f /usr/local/bin/npu-smi
53+maybe_mount f /etc/ascend_install.info
38# Unset => the eval's multi-card mode auto-detects every card, which is the normal full-run posture.54# Unset => the eval's multi-card mode auto-detects every card, which is the normal full-run posture.
39[[ -n "${ASCEND_RT_VISIBLE_DEVICES:-}" ]] && NPU_FLAGS+=(-e ASCEND_RT_VISIBLE_DEVICES="${ASCEND_RT_VISIBLE_DEVICES}")55[[ -n "${ASCEND_RT_VISIBLE_DEVICES:-}" ]] && NPU_FLAGS+=(-e ASCEND_RT_VISIBLE_DEVICES="${ASCEND_RT_VISIBLE_DEVICES}")
40 56 
@@ -3,7 +3,9 @@
3 3 
4Required checks (any failure -> non-zero exit):4Required checks (any failure -> non-zero exit):
5 [1] python / torch / torch_npu importable5 [1] python / torch / torch_npu importable
6- [2] torch_npu sees at least one NPU device6+ [2] torch_npu sees a device AND a bare H2D copy works -- the copy half is the real gate on a
7+ 950-class SoC, where it routes through aclnnInplaceCopy and so needs OPS_MODE=refonly
8+ (910B does a plain aclrtMemcpy and is happy with OPS_MODE=none)
7 [3] CANN compiler version.info readable9 [3] CANN compiler version.info readable
8 [4] cann_bench_utils importable -- the V3 anti-cheat warmup/cache-clean provider, a hard10 [4] cann_bench_utils importable -- the V3 anti-cheat warmup/cache-clean provider, a hard
9 dependency of every evaluation, baked in at build time11 dependency of every evaluation, baked in at build time
@@ -25,24 +27,35 @@ failed = []
25 27 
26# [1] versions28# [1] versions
27try:29try:
30+ import platform
31+ 
28 import torch32 import torch
29 import torch_npu33 import torch_npu
30 34 
31 py = ".".join(str(v) for v in sys.version_info[:3])35 py = ".".join(str(v) for v in sys.version_info[:3])
32- print(f"[OK] [1] python {py}, torch {torch.__version__}, torch_npu {torch_npu.__version__}")36+ print(
37+ f"[OK] [1] python {py}, torch {torch.__version__}, torch_npu {torch_npu.__version__}"
38+ f" ({platform.machine()})"
39+ )
33except Exception as e:40except Exception as e:
34 print(f"[FAIL] [1] import/version: {e}")41 print(f"[FAIL] [1] import/version: {e}")
35 failed.append(1)42 failed.append(1)
36 43 
37-# [2] device visible44+# [2] device visible AND usable
38try:45try:
46+ import torch
39 import torch_npu47 import torch_npu
40 48 
41 count = torch_npu.npu.device_count()49 count = torch_npu.npu.device_count()
42 assert count > 0, f"device_count = {count}"50 assert count > 0, f"device_count = {count}"
43- print(f"[OK] [2] torch_npu.npu.device_count() = {count}")51+ name = torch.npu.get_device_name(0)
52+ got = torch.arange(4, dtype=torch.float32).npu().cpu().tolist()
53+ assert got == [0.0, 1.0, 2.0, 3.0], got
54+ print(f"[OK] [2] {count} x {name}; bare H2D copy works")
44except Exception as e:55except Exception as e:
45- print(f"[FAIL] [2] torch_npu device_count: {e}")56+ print(f"[FAIL] [2] device / H2D copy: {e}")
57+ if "ERR01007" in str(e) or "aclnnInplaceCopy" in str(e):
58+ print(" ^ this SoC routes the copy through aclnn -- rebuild with OPS_MODE=refonly")
46 failed.append(2)59 failed.append(2)
47 60 
48# [3] CANN intact61# [3] CANN intact
@@ -82,17 +95,20 @@ except Exception as e:
82 failed.append(5)95 failed.append(5)
83 96 
84# [6] builtin availability -- diagnostic only. matmul is the canonical builtin the framework's own97# [6] builtin availability -- diagnostic only. matmul is the canonical builtin the framework's own
85-# warmup used to call before cann_bench_utils replaced it, so it is the right probe.98+# warmup used to call before cann_bench_utils replaced it, so it is the right probe. Expect it to FAIL:
86-ops_mode = os.environ.get("OPS_MODE", "none")99+# that is the anti-cheat working. The error differs by posture -- 500001 LazyInitAclops when libopapi
100+# is absent (none), 561103 "Parse dynamic kernel config fail" when it is present but the kernel
101+# binaries were stripped (refonly).
102+posture = f"OPS_MODE={os.environ.get('OPS_MODE', '?')} NPU_ARCH={os.environ.get('NPU_ARCH', '?')}"
87try:103try:
88 import torch104 import torch
89 105 
90 a = torch.randn(64, 64, device="npu:0")106 a = torch.randn(64, 64, device="npu:0")
91 (a @ a).cpu()107 (a @ a).cpu()
92- print(f"[INFO] [6] builtin aclnn ops CAN launch (OPS_MODE={ops_mode}) -- submissions could call them")108+ print(f"[WARN] [6] builtin aclnn ops CAN launch ({posture}) -- submissions could cheat by calling them")
93except Exception as e:109except Exception as e:
94- print(f"[INFO] [6] builtin aclnn ops cannot launch (OPS_MODE={ops_mode}): {type(e).__name__}")110+ print(f"[INFO] [6] builtin aclnn ops blocked ({posture}): {type(e).__name__}")
95- print(f" ^ expected for a 0-ops/refonly image; submissions must ship their own kernel. ({str(e)[:120]})")111+ print(f" ^ expected -- submissions must ship their own kernel. ({str(e)[:120]})")
96 112 
97# [7] optional Triton-Ascend113# [7] optional Triton-Ascend
98triton_ascend_version = os.environ.get("TRITON_ASCEND_VERSION", "").strip()114triton_ascend_version = os.environ.get("TRITON_ASCEND_VERSION", "").strip()