| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 3 个月前 | ||
| 2 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 3 个月前 |
NPUKernelBench
Ascend NPU kernel evaluation and benchmarking framework.
Evaluates custom NPU kernel solutions (PyTorch/torch_npu, Ascend C/CANN, Triton) for correctness and performance by comparing them against reference implementations.
Installation
git clone <repo-url>
cd npu-kernelbench
uv sync
Requires Python >=3.10 and Ascend NPU hardware with torch (CPU version) and torch-npu installed from Huawei's repository. The two packages must use the same version.
Quick Start
# Evaluate a single kernel (problem_dir mode)
uv run npu-kernelbench eval tests/samples/matmul_triton \
--solution tests/samples/matmul_triton/solution.json
# Evaluate with explicit paths
uv run npu-kernelbench eval \
--definition tests/samples/gelu/definition.json \
--workload tests/samples/gelu/workload.jsonl \
--solution path/to/solution.json
# Run a full dataset level
uv run npu-kernelbench run-dataset data/kernel_generator --level level1
# Run with aggregated HTML report
uv run npu-kernelbench run-dataset data/kernel_generator \
--level level1 --html report.html
# Run multiple levels with a limit
uv run npu-kernelbench run-dataset data/kernel_generator \
--level level1 --level level2 --limit 10
# Use isolated runner instead of persistent
uv run npu-kernelbench eval tests/samples/matmul_triton \
--solution solution.json --runner isolated
# Save traces to a specific file
uv run npu-kernelbench eval tests/samples/gelu \
--solution solution.json -o my_traces.jsonl
# Save logs to file
uv run npu-kernelbench eval tests/samples/gelu \
--solution solution.json --log-file eval.log --verbose
CLI Commands
| Command | Description |
|---|---|
eval |
Evaluate a single kernel solution |
run-dataset |
Batch evaluate all kernels in a dataset directory |
eval options
--definition PATH Path to definition.json
--workload PATH Path to workload.jsonl
--solution PATH Path to solution.json (required)
--config PATH Path to benchmark config JSON
--compile-timeout N Compilation timeout in seconds (default: 120)
--timeout N Evaluation subprocess timeout in seconds (default: 600)
--limit N Max workloads to evaluate
--runner MODE "persistent" (default) or "isolated"
--npu-ids IDS Comma-separated NPU device IDs (default: 0)
--profile-dir PATH Save profiling traces to this directory
--work-dir PATH Working directory for staging (default: /tmp/npu-kernelbench)
--log-file PATH Write logs to this file
--json Print trace JSON to stdout
--verbose, -v Enable debug logging
-o, --output PATH Write trace JSONL to this file (default: /tmp/npu-kernelbench/traces.jsonl)
--html PATH Generate HTML report at the given path
--resume Skip workloads/kernels already present in -o (resume after interruption)
run-dataset options
-l, --level TEXT Levels to evaluate (default: level1, repeatable)
--limit N Max workloads per kernel
--solution-dir PATH Directory containing solution files per kernel
--timeout N Evaluation subprocess timeout in seconds (default: 600)
--runner MODE "persistent" (default) or "isolated"
--npu-ids IDS Comma-separated NPU device IDs (default: 0)
--profile-dir PATH Save profiling traces to this directory
--work-dir PATH Working directory for staging (default: /tmp/npu-kernelbench)
--log-file PATH Write logs to this file
--verbose, -v Show verbose output
-o, --output PATH Write trace JSONL to this file (default: /tmp/npu-kernelbench/traces.jsonl)
--html PATH Generate aggregated HTML report at the given path
--resume Skip workloads/kernels already present in -o (resume after interruption)
Traces are always written to JSONL. Default path: /tmp/npu-kernelbench/traces.jsonl.
Use -o to specify a different path.
Generate HTML Report
# Generate HTML report alongside the default JSONL output
uv run npu-kernelbench eval tests/samples/matmul_triton \
--solution solution.json --html report.html
Open report.html in any browser to view the results.
Resume an Interrupted Run
# A batch run was interrupted — resume without redoing completed work.
# Workloads (and fully-done kernels) already present in -o are skipped;
# the output file keeps its existing results and appends only the new ones.
uv run npu-kernelbench run-dataset data/kernel_generator \
--level level1 -o runs.jsonl --resume
# Works the same way for eval, and for incremental runs where new kernels
# were added to the dataset — only the missing workloads run.
uv run npu-kernelbench eval tests/samples/matmul_triton \
--solution solution.json -o runs.jsonl --resume
Resume reads from the -o file (default /tmp/npu-kernelbench/traces.jsonl).
Reuse the same -o to continue a specific run; delete the file for a clean slate.
Configuration
See config.json for all options:
{
"warmup_runs": 5,
"iterations": 10,
"seed": 200,
"timeout": 600,
"compile_timeout": 120,
"benchmark_reference": true,
"npu_id": 0,
"npu_ids": [0],
"runner": "persistent",
"work_dir": "/tmp/npu-kernelbench",
"profile_dir": "",
"health_check_interval": 10,
"max_solution_failures": 3
}
Architecture
cli/main.py Click CLI (eval, run-dataset)
│
▼
runner/ Orchestration
├── runner.py Runner ABC
├── isolated_runner.py Process-per-evaluation (full NPU isolation)
├── persistent_runner.py Long-lived worker processes (default)
└── eval.py Standalone eval driver (staged to work_dir)
│
├── builder/ Transforms Solution → Runnable
│ ├── python_builder.py PyTorch, Triton, TileLang (importlib)
│ └── ascendc_builder.py Ascend C (CANN compilation)
│
├── evaluator/ Evaluation pipeline
│ ├── generator.py Input tensor generation (gen_inputs, _rand_tensor)
│ ├── correctness.py Numerical correctness checking
│ ├── performance.py NPU profiler-based timing
│ ├── anti_hack.py Reward hack detection
│ ├── evaluator.py Evaluator ABC + BaselineHandle
│ ├── default.py DefaultEvaluator (general-purpose kernels)
│ └── registry.py EvaluatorRegistry (priority-ordered resolution)
│
├── data/ Pydantic v2 models
│ ├── definition.py Definition (kernel spec, symbolic axes)
│ ├── workload.py Workload (InputSpec types, tolerances)
│ ├── solution.py Solution (source files, build spec)
│ ├── trace.py Trace, Evaluation, Correctness, Performance
│ ├── shapes.py Axis types (const, var, expr) + expression resolver
│ └── dtypes.py DType enum + torch/python dtype conversions
│
└── report/ Output formatting
├── formatter.py Rich table output
├── score.py Score computation
├── json_reporter.py JSON export
└── html_reporter.py HTML report generation
Execution flow
CLI → Runner pre_build (stage sources + eval files, compile)
→ Runner run_workload (per workload)
→ subprocess worker:
→ Builder loads Solution
→ Evaluator.build_baseline: gen_inputs → run reference
→ Evaluator.evaluate: run solution → correctness → performance → Trace
→ Report formats results (Rich table + score summary)
→ Traces saved to JSONL (default: work_dir/traces.jsonl)
Standalone reproduction
After evaluation, work_dir contains all files needed to reproduce:
work_dir/
├── definition.json
├── workload.jsonl
├── solution.json
├── eval.py ← cd work_dir && python eval.py
├── sources/
├── kernel.so (AscendC only)
└── profile/
Data Format
definition.json
Defines a kernel's interface: symbolic axes, tensor specs, and reference implementation.
{
"name": "matmul_triton",
"op_type": "matmul",
"axes": {
"M": {"type": "var", "description": "rows of A and C"},
"K": {"type": "var", "description": "inner dimension"},
"N": {"type": "var", "description": "cols of B and C"}
},
"inputs": {
"a": {"shape": ["M", "K"], "dtype": "float32"},
"b": {"shape": ["K", "N"], "dtype": "float32"}
},
"outputs": {
"c": {"shape": ["M", "N"], "dtype": "float32"}
},
"reference": "import torch\ndef run(a, b):\n return torch.mm(a, b)\n"
}
Axis types:
"const"— fixed value, e.g.{"type": "const", "value": 2048}"var"— determined by workload at runtime"expr"— computed from other axes, e.g.{"type": "expr", "expression": "M * 2"}
workload.jsonl (one JSON object per line)
Specifies concrete axis values and input data generation strategy.
{"axes": {"M": 64, "K": 128, "N": 64}, "inputs": {"a": {"type": "random", "dtype": "float32"}, "b": {"type": "random", "dtype": "float32"}}, "uuid": "mm-001"}
{"axes": {"M": 128, "K": 256, "N": 128}, "inputs": {"a": {"type": "random", "dtype": "float16"}, "b": {"type": "random", "dtype": "float16"}}, "uuid": "mm-002"}
{"axes": {"M": 128, "K": 256, "N": 128}, "inputs": {"a": {"type": "random", "dtype": "bfloat16"}, "b": {"type": "random", "dtype": "bfloat16"}}, "uuid": "mm-003"}
InputSpec types:
{"type": "random", "dtype": "float16"}— random tensor with optional dtype override{"type": "scalar", "value": 2.5}— scalar literal{"type": "custom"}— generated viacustom_inputs_entrypoint{"type": "safetensors", "path": "...", "tensor_key": "..."}— loaded from safetensors file
solution.json
A concrete kernel implementation.
{
"name": "my_matmul",
"definition": "matmul_triton",
"author": "developer",
"spec": {
"languages": ["triton"],
"target_hardware": ["ascend_910b"],
"entry_point": "kernel.py::matmul",
"destination_passing_style": false
},
"sources": [
{
"path": "kernel.py",
"content": "import torch\nimport triton\n...\ndef matmul(a, b):\n ...\n"
}
]
}
Supported languages: python, pytorch, triton, ascendc.
Supported hardware: ascend_910b, ascend_910c, LOCAL.
Sample Problems
The tests/samples/ directory contains reference problems:
| Directory | Description | Language |
|---|---|---|
gelu/ |
GELU activation function | — |
identity/ |
Identity function | — |
matmul_pytorch/ |
Matrix multiplication | PyTorch |
matmul_triton/ |
Matrix multiplication (Triton kernel) | Triton |
vector_add_ascendc/ |
Vector addition | Ascend C |
Testing
# Run all tests
uv run pytest tests/
# Run a single test file
uv run pytest tests/npu_kernelbench/data/test_definition.py -v
# Run tests matching a keyword
uv run pytest tests/ -k "test_correctness" -v
# Run integration tests (requires NPU)
uv run pytest tests/test_integration.py -v
Tests marked with requires_npu or ascend_c will skip when those
dependencies aren't available.
Build Wheel
pip install build
python -m build --wheel
Acknowledgements
NPUKernelBench's overall architecture and several implementation patterns (notably input tensor generation) draw inspiration from FlashInfer-Bench (docs), an open-source benchmark suite for LLM kernels released under the Apache 2.0 License.