Quick Start
This document provides two quick start scenarios to help you quickly get started with cluster scheduling powered by Ascend NPUs.
- 10-Minute Quick Start: Deploy only Ascend Device Plugin, use the Kubernetes native scheduler to schedule regular pods, and quickly verify NPU resource scheduling capabilities. This is suitable for beginners who want a quick experience.
- E2E Training Service Quick Start: Deploy all the cluster scheduling components (NodeD, Ascend Device Plugin, Ascend Docker Runtime, Volcano, ClusterD, and Ascend Operator). This scenario takes a PyTorch training job as an example to describe the E2E distributed training process.
You can choose the appropriate entry path based on your actual needs.
Environment Preparation
Ensure that the cluster environment has been set up.
-
Kubernetes has been installed on all nodes, with supported versions 1.17.x~1.34.x. If you need to install Volcano, install Kubernetes version 1.19.x or later. For specific Kubernetes versions, see the corresponding Kubernetes versions on the Volcano official website. To obtain the software package, see the Kubernetes community.
-
Docker has been installed on all nodes, with supported versions 18.09.x~28.5.1. To obtain the software package, see the Docker community or official website.
-
The corresponding firmware and drivers have been installed on all nodes. For the firmware and driver installation steps for the Atlas 800T A2 training server, see the Atlas A2 Center Inference and Training Hardware 26.0.RC1 NPU Driver and Firmware Installation Guide.
-
Check whether npu-smi and hccn_tool tools can run normally on the host.
Note
- Refer to the Ascend Training Solution Version Mapping to confirm whether the firmware and driver versions are compatible with the cluster scheduling components.
- The NPU driver and firmware versions can be queried using the
npu-smi info -t board -i NPU IDcommand. In the example output, theSoftware Versionfield indicates the NPU driver version, and theFirmware Versionfield indicates the NPU firmware version. - In the following text,
{xxx}takes the value910as the chip model.
10-Minute Quick Start
Overview
This tutorial guides you through setting up the most simplified Ascend NPU cluster scheduling environment in 10 minutes, using only:
- Ascend Device Plugin: Responsible for NPU device discovery and resource reporting.
- Kubernetes native scheduler: No additional scheduling components required.
- Regular pod: Responsible for verifying the NPU scheduling capability.
Environment Requirements
| Requirement | Description |
|---|---|
| Compute node | Atlas 800T A2 training server (Arm64) as an example |
| Driver version | Ascend driver matching the server |
Pre-check
Ensure that the NPU driver is correctly installed:
# Check NPU status. The chip information will be displayed in the command output
npu-smi info
Adding Labels to NPU Nodes
# Get the node name
kubectl get nodes
# Add necessary labels to NPU nodes (replace worker01 with the actual node name)
kubectl label nodes worker01 workerselector=dls-worker-node
kubectl label nodes worker01 accelerator=huawei-Ascend910
Deploying Ascend Device Plugin
1. Pull the Ascend Device Plugin image
# Pull the Ascend Device Plugin image from the Huawei Cloud image repository
docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-k8sdeviceplugin:v26.0.0
# Add a local label to the image
docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-k8sdeviceplugin:v26.0.0 ascend-k8sdeviceplugin:v26.0.0
2 Deploy Ascend Device Plugin
# Pull Configuration File
mkdir /tmp/devicePlugin
cd /tmp/devicePlugin
wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64.zip
unzip Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64.zip
# Deploy Ascend Device Plugin
kubectl apply -f device-plugin-910-v26.0.0.yaml
3 Verify the Deployment
# Check the Pod Status
kubectl get pod -n kube-system
# Expected Output
NAME READY STATUS RESTARTS AGE
...
ascend-device-plugin-daemonset-d5ctz 1/1 Running 0 11s
...
Verifying NPU Resources
# View the NPU resources of the node
kubectl describe node worker01 | grep -A 10 "huawei.com/Ascend910"
# Expected output (showing the number of available NPUs)
huawei.com/Ascend910: 8
huawei.com/Ascend910: 8
Scheduling an NPU Pod
1 Create a test pod configuration file
Create npu-test-pod.yaml:
apiVersion: v1
kind: Pod
metadata:
name: npu-test
spec:
nodeSelector:
workerselector: dls-worker-node
containers:
- name: npu-container
image: ubuntu:22.04
command: ["/bin/bash", "-c", "sleep 3600"]
resources:
limits:
huawei.com/Ascend910: 1 # Request 1 NPU
requests:
huawei.com/Ascend910: 1
volumeMounts:
- name: ascend-driver
mountPath: /usr/local/Ascend/driver
readOnly: true
volumes:
- name: ascend-driver
hostPath:
path: /usr/local/Ascend/driver
2 Deploy the test pod
kubectl apply -f npu-test-pod.yaml
3 Verify pod scheduling
# Check pod status
kubectl get pods npu-test -o wide
# Expected output (STATUS = Running indicates successful scheduling)
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE
npu-test 1/1 Running 0 10s 10.244.1.2 worker01 <none>
Verifying NPU Access
# Enter the container to verify NPU availability
kubectl exec -it npu-test -- /bin/bash
# Execute `npu-smi info` inside the container to check NPU information
export LD_LIBRARY_PATH=/usr/local/Ascend/driver/lib64/common:/usr/local/Ascend/driver/lib64/driver:${LD_LIBRARY_PATH}
npu-smi info
Cleaning Up Test Resources
# Delete the test pod
kubectl delete pod npu-test
# Delete Ascend Device Plugin (if needed)
kubectl delete -f device-plugin-910-v26.0.0.yaml
FAQs
| Issue | Cause | Solution |
|---|---|---|
| Pod remains in Pending state | Insufficient NPU resources or mismatched node labels | Check kubectl describe pod and node labels. |
| Ascend Device Plugin boot failure | Incorrect driver path | Check whether /usr/local/Ascend/driver exists. |
E2E Training Service Quick Start
This section uses two Atlas 800T A2 training servers (one as the management node and one as the compute node) as an example to guide developers through quickly installing NodeD, Ascend Device Plugin, Ascend Docker Runtime, Volcano, ClusterD, and Ascend Operator, and using the full-NPU scheduling feature to quickly submit a training job.
Procedure
Table 1 Key procedures
| Procedure | Description | For More Information |
|---|---|---|
| Installing Components | Using Atlas 800T A2 training servers as an example, this walks you through quickly installing cluster scheduling components on Ascend devices. | Installation and Deployment |
| Delivering a Training Job | Using a simple PyTorch training job as an example, this helps you quickly understand the workflow for submitting a training job. | Basic Scheduling |
Installing Components
The following uses an Atlas 800T A2 training server as an example. For detailed installation steps and parameter descriptions for all components, see Installation and Deployment.
-
Log in to the compute node or management node as the
rootuser and create the component installation directories.-
Run the following commands in sequence to create the installation directories on the compute node. The following directories are examples only.
mkdir /tmp/noded mkdir /tmp/devicePlugin mkdir /tmp/Ascend-docker-runtime -
Run the following commands in sequence to create the installation directories on the management node. The following directories are examples only.
mkdir /tmp/ascend-volcano mkdir /tmp/ascend-operator mkdir /tmp/clusterd
-
-
Download software packages with your desired architecture. The AArch64 architecture is used as an example.
-
Run the following commands in sequence to obtain the NodeD, Ascend Device Plugin, and Ascend Docker Runtime installation packages on the compute node and decompress them.
cd /tmp/noded wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-noded_26.0.0_linux-aarch64.zip unzip Ascend-mindxdl-noded_26.0.0_linux-aarch64.zip cd /tmp/devicePlugin wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64.zip unzip Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64.zip cd /tmp/Ascend-docker-runtime wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-docker-runtime_26.0.0_linux-aarch64.run -
Run the following commands in sequence on the management node to obtain the Volcano, ClusterD, and Ascend Operator installation packages.
cd /tmp/ascend-volcano wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-volcano_26.0.0_linux-aarch64.zip unzip Ascend-mindxdl-volcano_26.0.0_linux-aarch64.zip cd /tmp/ascend-operator wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-ascend-operator_26.0.0_linux-aarch64.zip unzip Ascend-mindxdl-ascend-operator_26.0.0_linux-aarch64.zip cd /tmp/clusterd wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-clusterd_26.0.0_linux-aarch64.zip unzip Ascend-mindxdl-clusterd_26.0.0_linux-aarch64.zip
-
-
Pull component images.
-
Run the following commands in sequence to pull the component images on the compute node.
cd /tmp/noded docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/noded:v26.0.0 docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/noded:v26.0.0 noded:v26.0.0 cd /tmp/devicePlugin docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-k8sdeviceplugin:v26.0.0 docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-k8sdeviceplugin:v26.0.0 ascend-k8sdeviceplugin:v26.0.0 -
Run the following commands in sequence to create component images on the management node.
cd /tmp/ascend-volcano/volcano-v1.7.0 docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/vc-scheduler:v1.7.0-v26.0.0 docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/vc-scheduler:v1.7.0-v26.0.0 volcanosh/vc-scheduler:v1.7.0-v26.0.0 docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/vc-controller-manager:v1.7.0-v26.0.0 docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/vc-controller-manager:v1.7.0-v26.0.0 volcanosh/vc-controller-manager:v1.7.0-v26.0.0 cd /tmp/ascend-operator docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-operator:v26.0.0 docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-operator:v26.0.0 ascend-operator:v26.0.0 cd /tmp/clusterd docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/clusterd:v26.0.0 docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/clusterd:v26.0.0 clusterd:v26.0.0
-
-
Create node labels.
Note
If the message "already has a value... and --overwrite is false" is displayed when you run the node label creation command, the label already exists. You can use
--overwriteto overwrite it.-
Run the following command on the Kubernetes management node to query the node name.
kubectl get nodeExample output:
NAME STATUS ROLES AGE VERSION worker01 Ready worker 23h v1.17.3 -
Run the following commands in sequence to create a node label for the compute node (
worker01as an example).kubectl label nodes worker01 node-role.kubernetes.io/worker=worker kubectl label nodes worker01 workerselector=dls-worker-node kubectl label nodes worker01 host-arch=huawei-arm kubectl label nodes worker01 accelerator=huawei-Ascend910 kubectl label nodes worker01 accelerator-type=module-{xxx}b-8 #Enter the chip model kubectl label nodes worker01 nodeDEnable=on -
Run the following command to create a node label for the management node (
master01as an example).kubectl label nodes master01 masterselector=dls-master-node
-
-
Create a user.
Note
Sudo privileges are required for user creation.
-
Run the following commands in sequence to create a username on the compute node.
useradd -d /home/hwMindX -u 9000 -m -s /usr/sbin/nologin hwMindX usermod -a -G HwHiAiUser hwMindX -
Run the following command to create a username on the management node.
useradd -d /home/hwMindX -u 9000 -m -s /usr/sbin/nologin hwMindX
-
-
Create log directories. Custom log directories are not supported.
Note
Sudo privileges are required for log directory creation.
-
Run the following commands in sequence to create log directories on the compute node.
mkdir -m 755 /var/log/mindx-dl chown root:root /var/log/mindx-dl mkdir -m 750 /var/log/mindx-dl/devicePlugin chown root:root /var/log/mindx-dl/devicePlugin mkdir -m 750 /var/log/mindx-dl/noded chown hwMindX:hwMindX /var/log/mindx-dl/noded -
Run the following commands in sequence to create log directories on the management node.
mkdir -m 755 /var/log/mindx-dl chown root:root /var/log/mindx-dl mkdir -m 750 /var/log/mindx-dl/volcano-controller chown hwMindX:hwMindX /var/log/mindx-dl/volcano-controller mkdir -m 750 /var/log/mindx-dl/volcano-scheduler chown hwMindX:hwMindX /var/log/mindx-dl/volcano-scheduler mkdir -m 750 /var/log/mindx-dl/ascend-operator chown hwMindX:hwMindX /var/log/mindx-dl/ascend-operator mkdir -m 750 /var/log/mindx-dl/clusterd chown hwMindX:hwMindX /var/log/mindx-dl/clusterd
-
-
Run the following command on any node to create the namespace.
kubectl create ns mindx-dl -
Install components.
-
Run the following commands in sequence to install Ascend Docker Runtime on the host of the compute node.
cd /tmp/Ascend-docker-runtime chmod u+x Ascend-docker-runtime_26.0.0_linux-aarch64.run ./Ascend-docker-runtime_26.0.0_linux-aarch64.run --install systemctl daemon-reload && systemctl restart docker -
On the compute node, run the following commands in sequence to install components.
cd /tmp/noded kubectl apply -f noded-v26.0.0.yaml cd /tmp/devicePlugin kubectl apply -f device-plugin-volcano-v26.0.0.yaml -
On the *management node, run the following commands in sequence to install components.
cd /tmp/ascend-operator kubectl apply -f ascend-operator-v26.0.0.yaml cd /tmp/ascend-volcano/volcano-v1.7.0 # If you are using Volcano version 1.9.0, change it to v1.9.0. kubectl apply -f volcano-v1.7.0.yaml cd /tmp/clusterd kubectl apply -f clusterd-v26.0.0.yaml -
Run the following command to check whether the components have started successfully.
kubectl get pod -ATaking ClusterD as an example, the example output is as follows.
Runningindicates that the component has started successfully.NAME READY STATUS RESTARTS AGE ... clusterd-fd6t8 1/1 Running 0 74s ...
-
Delivering a Training Job
-
Prepare an image.
Download the ascend-pytorch training image (24.0.X) from the Ascend Image Repository according to the system architecture (Arm/x86_64). Modify the training base image by changing the default user in the container to
root. The image does not contain training scripts, code, or other files. During training, files such as training scripts and code are typically mapped into the container using the mount method. -
Perform script adaptation.
-
Download "ResNet50_ID4149_for_PyTorch" from the master branch of the PyTorch Code Repository as the training code.
-
Prepare the dataset corresponding to ResNet-50 on your own, and comply with the corresponding specifications when using it.
-
The administrator uploads the dataset to the storage node. Go to the
/data/atlas_dls/publicdirectory and upload the dataset to any location, such as/data/atlas_dls/public/dataset/resnet50/imagenet.root@ubuntu:/data/atlas_dls/public/dataset/resnet50/imagenet# pwd -
Decompress the training code downloaded in Step 1 to the local machine, and upload the
ModelZoo-PyTorch/PyTorch/built-in/cv/classification/ResNet50_ID4149_for_PyTorchdirectory from the decompressed training code to the environment, for example, to the/data/atlas_dls/public/code/path. -
In the
/data/atlas_dls/public/code/ResNet50_ID4149_for_PyTorchpath, comment out the following code inmain.py.def main(): args = parser.parse_args() os.environ['MASTER_ADDR'] = args.addr #os.environ['MASTER_PORT'] = '29501' # Comment out this line of code if os.getenv('ALLOW_FP32', False) and os.getenv('ALLOW_HF32', False): raise RuntimeError('ALLOW_FP32 and ALLOW_HF32 cannot be set at the same time!') elif os.getenv('ALLOW_HF32', False): torch.npu.conv.allow_hf32 = True elif os.getenv('ALLOW_FP32', False): torch.npu.conv.allow_hf32 = False torch.npu.matmul.allow_hf32 = False -
Go to the mindcluster-deploy repository, switch to the corresponding version branch according to mindcluster-deploy Open-Source Repository Version Description, obtain the
train_start.shfile from thesamples/train/basic-training/without-ranktable/pytorchdirectory, and construct the following directory structure under the/data/atlas_dls/public/code/ResNet50_ID4149_for_PyTorch/scriptspath.root@ubuntu:/data/atlas_dls/public/code/ResNet50_ID4149_for_PyTorch/scripts# scripts/ ├── train_start.sh
-
-
Prepare the job YAML.
-
Go to the mindcluster-deploy repository. Based on the mindcluster-deploy Open-Source Repository Version Description, switch to the corresponding version branch and obtain the
pytorch_standalone_acjob_<i>\{xxx\}</i>.yamlfile from thesamples/train/basic-training/without-ranktable/pytorchdirectory ({xxx} indicates the chip model). The example defaults to a single-server single-device job. -
Modify the example YAML and upload it to any file path after modification. For detailed descriptions of each parameter in the following YAML, see Table 1.
apiVersion: mindxdl.gitee.com/v1 kind: AscendJob ... spec: ... replicaSpecs: Master: ... spec: nodeSelector: host-arch: huawei-arm accelerator-type: module-{xxx}b-8 # Change from the original `card-{xxx}b-2` to `module-{xxx}b-8`, where `{xxx}` indicates the chip model. containers: - name: ascend image: pytorch-test:latest # Modify to the image name obtained in Step 1. ... resources: limits: huawei.com/Ascend910: 1 requests: huawei.com/Ascend910: 1 ... volumes: - name: code nfs: #If the NFS service is not installed, change nfs to hostPath and delete server: 127.0.0.1 server: 127.0.0.1 path: "/data/atlas_dls/public/code/ResNet50_ID4149_for_PyTorch/" - name: data nfs: #If the NFS service is not installed, change nfs to hostPath and delete server: 127.0.0.1 server: 127.0.0.1 path: "/data/atlas_dls/public/dataset/" - name: output nfs: #If the NFS service is not installed, change nfs to hostPath and delete server: 127.0.0.1 server: 127.0.0.1 path: "/data/atlas_dls/output/" ...
-
-
Run the following command to deliver a single-server single-device job.
kubectl apply -f pytorch_standalone_acjob_{xxx}.yaml -
Run the following command to check the pod running status.
kubectl get pod --all-namespaces -o wideThe example output is as follows. If
Runningappears, the job is running normally.NAMESPACE NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES default default-test-pytorch-master-0 1/1 Running 0 6s 192.168.244.xxx worker01 <none> <none>Note
If the training job remains in the
Pendingstate after being delivered, refer to Training Job in Pending State, Cause: nodes are unavailable or Job in Pending State Due to Insufficient Resources for troubleshooting. -
View the training results.
-
Run the following command on any node to view the training results.
kubectl logs -n Namespace Pod nameFor example:
kubectl logs -n default default-test-pytorch-master-0 -
View the training log. If the following content appears, training succeeds.
[20251218-20:31:57] [MindXDL Service Log]server id is: 0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=7 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=6 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=5 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=4 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=3 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=2 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/bin/python /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py --data=/job/data/resnet50/imagenet --amp --arch=resnet50 --seed=49 -j=128 --world-size=1 --lr=1.6 --dist-backend=hccl --multiprocessing-distributed --epochs=1 --batch-size=512 --gpu=1 --multiprocessing-distributed --addr=10.106.227.104 --world-size=1 --rank=0 /usr/local/python3.10.5/lib/python3.10/site-packages/torchvision/io/image.py:13: UserWarning: Failed to load image Python extension: ''If you don't plan on using image functionality from `torchvision.io`, you can ignore this warning. Otherwise, there might be something wrong with your environment. Did you have `libjpeg` or `libpng` installed before building `torchvision` from source? warn( [2025-12-18 20:32:02] [WARNING] [470] profiler.py: Invalid parameter export_type: None, reset it to text. /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py:201: UserWarning: You have chosen to seed training. This will turn on the CUDNN deterministic setting, which can slow down your training considerably! You may see unexpected behavior when restarting from checkpoints. warnings.warn('You have chosen to seed training. ' /job/code/No_Rank_ResNet50_ID4149_for_PyTorch/main.py:208: UserWarning: You have chosen a specific GPU. This will completely disable data parallelism. warnings.warn('You have chosen a specific GPU. This will completely ' Use GPU: 0 for training => creating model 'resnet50'
-