已合并
chore(precision-test): 同步 test_selector 上游并适配 torch_npu 改为 pytorch (#18) #44
tzing_t创建于 5 天前
chore(precision-test): 同步 test_selector 上游并适配 torch_npu 改为 pytorch (#18) #44
已合并
共 21 个文件变更+595-180
| @@ -1,12 +1,12 @@ | |||
| 1 | # precision-test | 1 | # precision-test |
| 2 | 2 | ||
| 3 | -<!-- 基于代码覆盖率变更的精准测试用例选择器——分析 PR 代码变更,按覆盖率数据以行、函数两种粒度筛选相关测试用例(文件级作为重命名/删除场景的内部兜底),避免每次 PR 跑全量用例。统一支持 vllm_ascend / sglang / torch_npu 三仓库。 --> | 3 | +<!-- 基于代码覆盖率变更的精准测试用例选择器——分析 PR 代码变更,按覆盖率数据以行、函数两种粒度筛选相关测试用例(文件级作为重命名/删除场景的内部兜底),避免每次 PR 跑全量用例。统一支持 vllm_ascend / sglang / pytorch 三仓库。 --> |
| 4 | 4 | ||
| 5 | ## 核心特性 (Features) | 5 | ## 核心特性 (Features) |
| 6 | 6 | ||
| 7 | -- **三仓库统一**:一套插件支持 `vllm_ascend` / `sglang` / `torch_npu`,通过 `repo-name` 选择仓库适配器 | 7 | +- **三仓库统一**:一套插件支持 `vllm_ascend` / `sglang` / `pytorch`,通过 `repo-name` 选择仓库适配器 |
| 8 | - **多粒度匹配**:行级 + 函数级并行匹配,结果合并去重;文件级匹配作为重命名/删除场景的内部兜底(不可手动开关) | 8 | - **多粒度匹配**:行级 + 函数级并行匹配,结果合并去重;文件级匹配作为重命名/删除场景的内部兜底(不可手动开关) |
| 9 | -- **PR 源自动适配**:vllm_ascend / sglang 走 GitHub(`github-pr`),torch_npu 走 GitCode(`gitcode-pr`),两者互斥 | 9 | +- **PR 源自动适配**:三仓库均走 GitHub(`github-pr`,pytorch 检测上游 pytorch/pytorch);`gitcode-pr` 作为 GitCode PR 通用入口,与 `github-pr` 互斥 |
| 10 | - **基于覆盖率选择**:利用覆盖率数据(SQLite 格式)精准识别受代码变更影响的测试用例 | 10 | - **基于覆盖率选择**:利用覆盖率数据(SQLite 格式)精准识别受代码变更影响的测试用例 |
| 11 | - **安全加固**:内置 DoS 防护、子进程环境变量白名单、路径与 PR 标识校验 | 11 | - **安全加固**:内置 DoS 防护、子进程环境变量白名单、路径与 PR 标识校验 |
| 12 | - **可配置**:匹配粒度、阈值、虚拟环境等均可通过 action 输入参数调整 | 12 | - **可配置**:匹配粒度、阈值、虚拟环境等均可通过 action 输入参数调整 |
| @@ -57,7 +57,7 @@ jobs: | |||
| 57 | fi | 57 | fi |
| 58 | ``` | 58 | ``` |
| 59 | 59 | ||
| 60 | -#### torch_npu(GitCode PR) | 60 | +#### pytorch(GitHub PR) |
| 61 | 61 | ||
| 62 | ```yaml | 62 | ```yaml |
| 63 | name: 精准测试 | 63 | name: 精准测试 |
| @@ -79,8 +79,8 @@ jobs: | |||
| 79 | - name: 精准测试选择 | 79 | - name: 精准测试选择 |
| 80 | uses: openlibing/test-action/precision-test@v2.0.0 | 80 | uses: openlibing/test-action/precision-test@v2.0.0 |
| 81 | with: | 81 | with: |
| 82 | - repo-name: torch_npu | 82 | + repo-name: pytorch |
| 83 | - gitcode-pr: ${{ atomgit.repository }}#${{ atomgit.event.pull_request.number }} | 83 | + github-pr: pytorch/pytorch#${{ github.event.pull_request.number }} |
| 84 | source-dir: covstub | 84 | source-dir: covstub |
| 85 | min-affected: 1 | 85 | min-affected: 1 |
| 86 | 86 | ||
| @@ -110,23 +110,23 @@ jobs: | |||
| 110 | 110 | ||
| 111 | ## 输入参数 (Inputs) | 111 | ## 输入参数 (Inputs) |
| 112 | 112 | ||
| 113 | -| 参数名 | 说明 | 必填 | 默认值 | | 113 | +| 参数名 | 说明 | 必填 | 默认值 | |
| 114 | -| ------------------------ | ----------------------------------------------------------------------------------------------------- | ------ | -------------------- | | 114 | +| ------------------------ | --------------------------------------------------------------------------------------------------------------- | ------ | -------------------- | |
| 115 | -| `github-pr` | GitHub PR(vllm_ascend / sglang),格式 `owner/repo#pr_number` 或仅 `pr_number`;与 `gitcode-pr` 互斥 | 二选一 | - | | 115 | +| `github-pr` | GitHub PR(vllm_ascend / sglang / pytorch),格式 `owner/repo#pr_number` 或仅 `pr_number`;与 `gitcode-pr` 互斥 | 二选一 | - | |
| 116 | -| `gitcode-pr` | GitCode PR(torch_npu),格式 `owner/repo#pr_number` 或仅 `pr_number`;与 `github-pr` 互斥 | 二选一 | - | | 116 | +| `gitcode-pr` | GitCode PR,格式 `owner/repo#pr_number` 或仅 `pr_number`;与 `github-pr` 互斥 | 二选一 | - | |
| 117 | -| `repo-name` | 仓库适配器:`vllm_ascend` / `sglang` / `torch_npu` | 否 | `vllm_ascend` | | 117 | +| `repo-name` | 仓库适配器:`vllm_ascend` / `sglang` / `pytorch` | 否 | `vllm_ascend` | |
| 118 | -| `source-dir` | 源代码目录(函数级匹配与噪音过滤需要) | 否 | `covstub` | | 118 | +| `source-dir` | 源代码目录(函数级匹配与噪音过滤需要) | 否 | `covstub` | |
| 119 | -| `map-file` | 测试用例映射文件路径 | 否 | `test_case_map.json` | | 119 | +| `map-file` | 测试用例映射文件路径 | 否 | `test_case_map.json` | |
| 120 | -| `coverage-dir` | 覆盖率数据目录;构建 map 时(`build-map: true` 或 map 文件不存在时)必填 | 否 | `coverage` | | 120 | +| `coverage-dir` | 覆盖率数据目录;构建 map 时(`build-map: true` 或 map 文件不存在时)必填 | 否 | `coverage` | |
| 121 | -| `build-map` | 是否强制重建测试用例映射 | 否 | `false` | | 121 | +| `build-map` | 是否强制重建测试用例映射 | 否 | `false` | |
| 122 | -| `min-affected` | 最小受影响行数阈值 | 否 | `1` | | 122 | +| `min-affected` | 最小受影响行数阈值 | 否 | `1` | |
| 123 | -| `dedup` | 是否启用去重(相同覆盖行仅保留一个用例) | 否 | `false` | | 123 | +| `dedup` | 是否启用去重(相同覆盖行仅保留一个用例) | 否 | `false` | |
| 124 | -| `enable-line-match` | 是否启用行级匹配 | 否 | `true` | | 124 | +| `enable-line-match` | 是否启用行级匹配 | 否 | `true` | |
| 125 | -| `disable-line-match` | 关闭行级匹配(**优先级高于 `enable-line-match`**) | 否 | `false` | | 125 | +| `disable-line-match` | 关闭行级匹配(**优先级高于 `enable-line-match`**) | 否 | `false` | |
| 126 | -| `enable-function-match` | 是否启用函数级匹配 | 否 | `true` | | 126 | +| `enable-function-match` | 是否启用函数级匹配 | 否 | `true` | |
| 127 | -| `disable-function-match` | 关闭函数级匹配(**优先级高于 `enable-function-match`**) | 否 | `false` | | 127 | +| `disable-function-match` | 关闭函数级匹配(**优先级高于 `enable-function-match`**) | 否 | `false` | |
| 128 | -| `skip-imports` | 函数级匹配时是否跳过 import 语句行 | 否 | `false` | | 128 | +| `skip-imports` | 函数级匹配时是否跳过 import 语句行 | 否 | `false` | |
| 129 | -| `skip-venv` | 跳过虚拟环境创建,直接使用系统 `python3` / `pip3` | 否 | `true` | | 129 | +| `skip-venv` | 跳过虚拟环境创建,直接使用系统 `python3` / `pip3` | 否 | `true` | |
| 130 | 130 | ||
| 131 | **匹配粒度规则**:`disable-*` 优先于 `enable-*`。即 `disable-line-match: true` 时无论 `enable-line-match` 取值如何,行级匹配均关闭。文件级匹配仅作为重命名/删除场景的内部兜底,不对外暴露开关。 | 131 | **匹配粒度规则**:`disable-*` 优先于 `enable-*`。即 `disable-line-match: true` 时无论 `enable-line-match` 取值如何,行级匹配均关闭。文件级匹配仅作为重命名/删除场景的内部兜底,不对外暴露开关。 |
| 132 | 132 | ||
| @@ -148,10 +148,10 @@ permissions: | |||
| 148 | 148 | ||
| 149 | ## 仓库差异对照 | 149 | ## 仓库差异对照 |
| 150 | 150 | ||
| 151 | -| 维度 | vllm_ascend | sglang | torch_npu | | 151 | +| 维度 | vllm_ascend | sglang | pytorch | |
| 152 | | ------------ | ------------------------- | ------------------------------------ | ------------------------------- | | 152 | | ------------ | ------------------------- | ------------------------------------ | ------------------------------- | |
| 153 | -| PR 源 | GitHub(`github-pr`) | GitHub(`github-pr`) | GitCode(`gitcode-pr`) | | 153 | +| PR 源 | GitHub(`github-pr`) | GitHub(`github-pr`) | GitHub(`github-pr`) | |
| 154 | -| 产品代码前缀 | `vllm_ascend/` | `python/sglang/` | `torch_npu/` | | 154 | +| 产品代码前缀 | `vllm_ascend/` | `python/sglang/` | `torch/` | |
| 155 | | 全量触发变更 | `csrc/` 目录(非 `.md`) | 无 | 无 | | 155 | | 全量触发变更 | `csrc/` 目录(非 `.md`) | 无 | 无 | |
| 156 | | 测试目录识别 | `tests__` 前缀 / `cpu-ut` | `____w__sglang__sglang__test__` 前缀 | 目录名含 `__` 或以 `test_` 开头 | | 156 | | 测试目录识别 | `tests__` 前缀 / `cpu-ut` | `____w__sglang__sglang__test__` 前缀 | 目录名含 `__` 或以 `test_` 开头 | |
| 157 | 157 | ||
| @@ -191,7 +191,7 @@ permissions: | |||
| 191 | | --------------- | ------ | ------------------------------------------------------------------- | | 191 | | --------------- | ------ | ------------------------------------------------------------------- | |
| 192 | | `GITHUB_TOKEN` | - | GitHub API token(vllm_ascend / sglang 拉 PR diff,可选,提升速率) | | 192 | | `GITHUB_TOKEN` | - | GitHub API token(vllm_ascend / sglang 拉 PR diff,可选,提升速率) | |
| 193 | | `GH_TOKEN` | - | GitHub API token 备选(`GITHUB_TOKEN` 未设置时回退) | | 193 | | `GH_TOKEN` | - | GitHub API token 备选(`GITHUB_TOKEN` 未设置时回退) | |
| 194 | -| `GITCODE_TOKEN` | - | GitCode API token(torch_npu 拉 PR diff,可选,提升速率) | | 194 | +| `GITCODE_TOKEN` | - | GitCode API token(`gitcode-pr` 拉 PR diff,可选,提升速率) | |
| 195 | 195 | ||
| 196 | > 插件通过子进程环境变量白名单(`SUBPROCESS_ENV_SAFE_NAMES`)向 Python 子进程显式传递 `PATH`、`HOME`、代理、SSL 证书、`GITHUB_TOKEN` / `GH_TOKEN` / `GITCODE_TOKEN` 等安全变量,屏蔽 `ACTIONS_*` / `ATOMGIT_*` / `INPUT_*` / `HUAWEICLOUD_*` 等平台敏感前缀,避免密钥泄露到 Python 进程。 | 196 | > 插件通过子进程环境变量白名单(`SUBPROCESS_ENV_SAFE_NAMES`)向 Python 子进程显式传递 `PATH`、`HOME`、代理、SSL 证书、`GITHUB_TOKEN` / `GH_TOKEN` / `GITCODE_TOKEN` 等安全变量,屏蔽 `ACTIONS_*` / `ATOMGIT_*` / `INPUT_*` / `HUAWEICLOUD_*` 等平台敏感前缀,避免密钥泄露到 Python 进程。 |
| 197 | 197 | ||
| @@ -223,8 +223,8 @@ GitHub / GitCode API 均需直连外网,内网环境必须配置代理: | |||
| 223 | PYTHON_EXEC_TIMEOUT_SECONDS: 900 | 223 | PYTHON_EXEC_TIMEOUT_SECONDS: 900 |
| 224 | GITCODE_TOKEN: ${{ secrets.GITCODE_TOKEN }} | 224 | GITCODE_TOKEN: ${{ secrets.GITCODE_TOKEN }} |
| 225 | with: | 225 | with: |
| 226 | - repo-name: torch_npu | 226 | + repo-name: pytorch |
| 227 | - gitcode-pr: ${{ atomgit.repository }}#${{ atomgit.event.pull_request.number }} | 227 | + github-pr: pytorch/pytorch#${{ github.event.pull_request.number }} |
| 228 | ``` | 228 | ``` |
| 229 | 229 | ||
| 230 | ## 数据格式 | 230 | ## 数据格式 |
| @@ -234,7 +234,7 @@ GitHub / GitCode API 均需直连外网,内网环境必须配置代理: | |||
| 234 | - **file**:文件路径信息 | 234 | - **file**:文件路径信息 |
| 235 | - **arc**:覆盖弧数据,含 `fromno`、`tono` 字段 | 235 | - **arc**:覆盖弧数据,含 `fromno`、`tono` 字段 |
| 236 | 236 | ||
| 237 | -覆盖率目录布局自动探测:测试目录下存在 `covdata/` 子目录(vllm_ascend / torch_npu)则从 `covdata/` 读取 `coverage.*`;无 `covdata/` 子目录(sglang)则直接读取测试目录下的 `coverage.*`。 | 237 | +覆盖率目录布局自动探测:测试目录下存在 `covdata/` 子目录(vllm_ascend / pytorch)则从 `covdata/` 读取 `coverage.*`;无 `covdata/` 子目录(sglang)则直接读取测试目录下的 `coverage.*`。 |
| 238 | 238 | ||
| 239 | ## 上下游串联 | 239 | ## 上下游串联 |
| 240 | 240 | ||
| @@ -1,16 +1,16 @@ | |||
| 1 | name: "Precision Test Selector" | 1 | name: "Precision Test Selector" |
| 2 | -description: "Coverage-based precision test selector with line, function, and file granularity (vllm_ascend / sglang / torch_npu)" | 2 | +description: "Coverage-based precision test selector with line, function, and file granularity (vllm_ascend / sglang / pytorch)" |
| 3 | author: "openlibing" | 3 | author: "openlibing" |
| 4 | 4 | ||
| 5 | inputs: | 5 | inputs: |
| 6 | github-pr: | 6 | github-pr: |
| 7 | - description: "GitHub PR (vllm_ascend / sglang), format: owner/repo#pr_number or just pr_number" | 7 | + description: "GitHub PR (vllm_ascend / sglang / pytorch), format: owner/repo#pr_number or just pr_number" |
| 8 | required: false | 8 | required: false |
| 9 | gitcode-pr: | 9 | gitcode-pr: |
| 10 | - description: "GitCode PR (torch_npu), format: owner/repo#pr_number or just pr_number; mutually exclusive with github-pr" | 10 | + description: "GitCode PR, format: owner/repo#pr_number or just pr_number; mutually exclusive with github-pr" |
| 11 | required: false | 11 | required: false |
| 12 | repo-name: | 12 | repo-name: |
| 13 | - description: "Repository adapter: vllm_ascend / sglang / torch_npu (default: vllm_ascend)" | 13 | + description: "Repository adapter: vllm_ascend / sglang / pytorch (default: vllm_ascend)" |
| 14 | required: false | 14 | required: false |
| 15 | default: "vllm_ascend" | 15 | default: "vllm_ascend" |
| 16 | source-dir: | 16 | source-dir: |
| @@ -25651,7 +25651,7 @@ const path = __nccwpck_require__(6928); | |||
| 25651 | /** | 25651 | /** |
| 25652 | * Available repository adapters supported by the precision-test tool. | 25652 | * Available repository adapters supported by the precision-test tool. |
| 25653 | */ | 25653 | */ |
| 25654 | -const AVAILABLE_REPOS = ["vllm_ascend", "sglang", "torch_npu"]; | 25654 | +const AVAILABLE_REPOS = ["vllm_ascend", "sglang", "pytorch"]; |
| 25655 | 25655 | ||
| 25656 | /** | 25656 | /** |
| 25657 | * Whitelist of environment variable names safe to pass to the Python subprocess. | 25657 | * Whitelist of environment variable names safe to pass to the Python subprocess. |
| @@ -1,10 +1,10 @@ | |||
| 1 | -"""test_selector - 覆盖率驱动的精准测试选择器(vllm_ascend / sglang / torch_npu 统一架构)。 | 1 | +"""test_selector - 覆盖率驱动的精准测试选择器(vllm_ascend / sglang / pytorch 统一架构)。 |
| 2 | 2 | ||
| 3 | 公共逻辑集中在 test_selector/ 包(diff 解析、噪音过滤、覆盖率选择、CLI 等), | 3 | 公共逻辑集中在 test_selector/ 包(diff 解析、噪音过滤、覆盖率选择、CLI 等), |
| 4 | 仓库差异通过 test_selector.repos.RepoAdapter 注入,保证架构统一、差异按仓库归类。 | 4 | 仓库差异通过 test_selector.repos.RepoAdapter 注入,保证架构统一、差异按仓库归类。 |
| 5 | 5 | ||
| 6 | 入口: | 6 | 入口: |
| 7 | - python -m test_selector --repo <vllm_ascend|sglang|torch_npu> ... | 7 | + python -m test_selector --repo <vllm_ascend|sglang|pytorch> ... |
| 8 | """ | 8 | """ |
| 9 | 9 | ||
| 10 | from .repos import AVAILABLE_REPOS, RepoAdapter, get_adapter | 10 | from .repos import AVAILABLE_REPOS, RepoAdapter, get_adapter |
| @@ -106,6 +106,13 @@ class CodeChangeDetector: | |||
| 106 | NamedTuple base) or module level, unless the annotation expression | 106 | NamedTuple base) or module level, unless the annotation expression |
| 107 | may have runtime side effects (calls, Annotated metadata) and the | 107 | may have runtime side effects (calls, Annotated metadata) and the |
| 108 | file has no ``from __future__ import annotations``. | 108 | file has no ``from __future__ import annotations``. |
| 109 | + - First-time bindings are excluded for pure insertions (needs base | ||
| 110 | + content): every added line is an assignment binding simple names | ||
| 111 | + that occur nowhere in the base file, with inert right-hand sides; | ||
| 112 | + no meaningful line ends in a comma (such lines are call/signature | ||
| 113 | + arguments, not statements); class-scope insertions follow dataclass | ||
| 114 | + field-order rules (plain assigns are not fields, annotated fields | ||
| 115 | + must land at or after the last existing field). | ||
| 109 | - Isolated blank-line deletion (neighbours not deleted): treated as a | 116 | - Isolated blank-line deletion (neighbours not deleted): treated as a |
| 110 | one-line insertion -> candidate pair (line above, line below). | 117 | one-line insertion -> candidate pair (line above, line below). |
| 111 | - Pure insertions and blank-deletion pairs are classified via ast of | 118 | - Pure insertions and blank-deletion pairs are classified via ast of |
| @@ -10,11 +10,11 @@ | |||
| 10 | 10 | ||
| 11 | 用法: | 11 | 用法: |
| 12 | python -m test_selector --repo vllm_ascend --github-pr owner/repo#123 | 12 | python -m test_selector --repo vllm_ascend --github-pr owner/repo#123 |
| 13 | - python -m test_selector --repo torch_npu --gitcode-pr Ascend/pytorch#123 | 13 | + python -m test_selector --repo torch_npu --github-pr pytorch/pytorch#123 |
| 14 | python -m test_selector --repo sglang --build-map --coverage-dir coverage | 14 | python -m test_selector --repo sglang --build-map --coverage-dir coverage |
| 15 | python vllm/test_selector.py --github-pr 123 # 薄入口(默认 vllm_ascend) | 15 | python vllm/test_selector.py --github-pr 123 # 薄入口(默认 vllm_ascend) |
| 16 | python sglang/test_selector.py --github-pr 123 # 薄入口(默认 sglang) | 16 | python sglang/test_selector.py --github-pr 123 # 薄入口(默认 sglang) |
| 17 | - python PyTorch/test_selector.py --gitcode-pr 123 # 薄入口(默认 torch_npu) | 17 | + python PyTorch/test_selector.py --github-pr 123 # 薄入口(默认 torch_npu,检测上游 pytorch/pytorch) |
| 18 | """ | 18 | """ |
| 19 | 19 | ||
| 20 | import argparse | 20 | import argparse |
| @@ -49,20 +49,21 @@ class CoverageSelector: | |||
| 49 | if not self.coverage_data_dir or not self.coverage_data_dir.exists(): | 49 | if not self.coverage_data_dir or not self.coverage_data_dir.exists(): |
| 50 | print(f" Warning: Coverage data directory not found: {self.coverage_data_dir}") | 50 | print(f" Warning: Coverage data directory not found: {self.coverage_data_dir}") |
| 51 | return test_cases | 51 | return test_cases |
| 52 | - for item in self.coverage_data_dir.iterdir(): | 52 | + for item in self.coverage_data_dir.rglob("*"): |
| 53 | if not item.is_dir(): | 53 | if not item.is_dir(): |
| 54 | continue | 54 | continue |
| 55 | name = item.name | 55 | name = item.name |
| 56 | # 目录名规则由仓库适配器决定(vllm: tests__ 前缀 / cpu-ut;sglang: ____w__... 前缀) | 56 | # 目录名规则由仓库适配器决定(vllm: tests__ 前缀 / cpu-ut;sglang: ____w__... 前缀) |
| 57 | if not self.adapter.is_test_case_dir(name): | 57 | if not self.adapter.is_test_case_dir(name): |
| 58 | continue | 58 | continue |
| 59 | - # 布局探测:covdata/ 子目录(vllm/torch_npu)或目录下直接放置(sglang) | 59 | + # 布局探测:covdata/ 子目录(vllm/pytorch)或目录下直接放置(sglang) |
| 60 | covdata_dir = item / "covdata" | 60 | covdata_dir = item / "covdata" |
| 61 | has_cov_files = any(item.glob(self.adapter.coverage_file_glob)) | 61 | has_cov_files = any(item.glob(self.adapter.coverage_file_glob)) |
| 62 | if covdata_dir.exists() and any(covdata_dir.glob(self.adapter.coverage_file_glob)): | 62 | if covdata_dir.exists() and any(covdata_dir.glob(self.adapter.coverage_file_glob)): |
| 63 | has_cov_files = True | 63 | has_cov_files = True |
| 64 | if has_cov_files: | 64 | if has_cov_files: |
| 65 | - test_cases.append(name) | 65 | + rel = item.relative_to(self.coverage_data_dir).as_posix() |
| 66 | + test_cases.append(rel) | ||
| 66 | return sorted(test_cases) | 67 | return sorted(test_cases) |
| 67 | 68 | ||
| 68 | def normalize_test_name(self, test_name: str) -> str: | 69 | def normalize_test_name(self, test_name: str) -> str: |
| @@ -142,7 +143,7 @@ class CoverageSelector: | |||
| 142 | 143 | ||
| 143 | file_lines_map = defaultdict(set) # filepath -> set of lines | 144 | file_lines_map = defaultdict(set) # filepath -> set of lines |
| 144 | 145 | ||
| 145 | - # 布局探测:vllm/torch_npu 将 coverage 放在 covdata/ 子目录,sglang 直接放在测试目录下 | 146 | + # 布局探测:vllm/pytroch 将 coverage 放在 covdata/ 子目录,sglang 直接放在测试目录下 |
| 146 | cov_dirs = [covdata_dir] if covdata_dir.exists() else [test_case_dir] | 147 | cov_dirs = [covdata_dir] if covdata_dir.exists() else [test_case_dir] |
| 147 | 148 | ||
| 148 | for cov_dir in cov_dirs: | 149 | for cov_dir in cov_dirs: |
| @@ -4,6 +4,8 @@ | |||
| 4 | 代码变更的场景: | 4 | 代码变更的场景: |
| 5 | - 纯注释/docstring 变更 | 5 | - 纯注释/docstring 变更 |
| 6 | - 纯类型注解变更(插入/替换/删除三个方向,可证明惰性时豁免) | 6 | - 纯类型注解变更(插入/替换/删除三个方向,可证明惰性时豁免) |
| 7 | +- 首次新增的变量绑定(纯插入块全部为绑定全新名字的带值赋值,RHS 惰性, | ||
| 8 | + 无尾逗号行——尾逗号行是调用/签名参数而非语句) | ||
| 7 | - 函数/类定义之间的空行插入 | 9 | - 函数/类定义之间的空行插入 |
| 8 | - 新增函数/类整体 | 10 | - 新增函数/类整体 |
| 9 | 11 | ||
| @@ -318,14 +320,27 @@ def _class_consumes_annotations(cls_node: ast.ClassDef, class_map: dict) -> bool | |||
| 318 | return False | 320 | return False |
| 319 | 321 | ||
| 320 | 322 | ||
| 321 | -def _collect_class_info(source: str) -> tuple[list[tuple], dict[str, ast.ClassDef], bool]: | 323 | +def _arg_names(args: ast.arguments) -> set[str]: |
| 322 | - """Parse source and return (scopes, class_map, has_future): | 324 | + """Names bound by a function/lambda parameter list.""" |
| 325 | + names = {a.arg for a in args.posonlyargs + args.args + args.kwonlyargs} | ||
| 326 | + if args.vararg: | ||
| 327 | + names.add(args.vararg.arg) | ||
| 328 | + if args.kwarg: | ||
| 329 | + names.add(args.kwarg.arg) | ||
| 330 | + return names | ||
| 323 | 331 | ||
| 324 | - scopes : [(lineno, end_lineno, col_offset, node)] of every | 332 | + |
| 325 | - function/class definition, for annotation scope resolution; | 333 | +def _collect_class_info(source: str) -> tuple[list[tuple], dict[str, ast.ClassDef], bool, set[str]]: |
| 326 | - class_map : {class name -> ClassDef} for same-file base resolution | 334 | + """Parse source and return (scopes, class_map, has_future, known_names): |
| 327 | - (a later definition shadows an earlier one); | 335 | + |
| 328 | - has_future: the file has ``from __future__ import annotations``. | 336 | + scopes : [(lineno, end_lineno, col_offset, node)] of every |
| 337 | + function/class definition, for annotation scope resolution; | ||
| 338 | + class_map : {class name -> ClassDef} for same-file base resolution | ||
| 339 | + (a later definition shadows an earlier one); | ||
| 340 | + has_future : the file has ``from __future__ import annotations``. | ||
| 341 | + known_names : every name the base file already binds or reads (Name | ||
| 342 | + occurrences, def/class names, import aliases, parameters), | ||
| 343 | + for first-time-binding detection. | ||
| 329 | """ | 344 | """ |
| 330 | tree = ast.parse(source) | 345 | tree = ast.parse(source) |
| 331 | scopes = [ | 346 | scopes = [ |
| @@ -344,7 +359,24 @@ def _collect_class_info(source: str) -> tuple[list[tuple], dict[str, ast.ClassDe | |||
| 344 | and any(alias.name == "annotations" for alias in n.names) | 359 | and any(alias.name == "annotations" for alias in n.names) |
| 345 | for n in tree.body | 360 | for n in tree.body |
| 346 | ) | 361 | ) |
| 347 | - return scopes, class_map, has_future | 362 | + known_names = set() |
| 363 | + for node in ast.walk(tree): | ||
| 364 | + if isinstance(node, ast.Name): | ||
| 365 | + known_names.add(node.id) | ||
| 366 | + elif isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)): | ||
| 367 | + known_names.add(node.name) | ||
| 368 | + known_names.update(_arg_names(node.args)) | ||
| 369 | + elif isinstance(node, ast.ClassDef): | ||
| 370 | + known_names.add(node.name) | ||
| 371 | + elif isinstance(node, ast.Lambda): | ||
| 372 | + known_names.update(_arg_names(node.args)) | ||
| 373 | + elif isinstance(node, ast.ImportFrom): | ||
| 374 | + for alias in node.names: | ||
| 375 | + known_names.add(alias.asname or alias.name) | ||
| 376 | + elif isinstance(node, ast.Import): | ||
| 377 | + for alias in node.names: | ||
| 378 | + known_names.add(alias.asname or alias.name.split(".")[0]) | ||
| 379 | + return scopes, class_map, has_future, known_names | ||
| 348 | 380 | ||
| 349 | 381 | ||
| 350 | def _insertion_scope(scopes: list[tuple], a: int, b: int, add_indent: int | None): | 382 | def _insertion_scope(scopes: list[tuple], a: int, b: int, add_indent: int | None): |
| @@ -391,7 +423,7 @@ def _pure_annotation_insert_reason( | |||
| 391 | """ | 423 | """ |
| 392 | if ann_ctx is None or not add_texts: | 424 | if ann_ctx is None or not add_texts: |
| 393 | return None | 425 | return None |
| 394 | - scopes, class_map, has_future = ann_ctx | 426 | + scopes, class_map, has_future, _known = ann_ctx |
| 395 | anns = _parse_pure_annotations(add_texts) | 427 | anns = _parse_pure_annotations(add_texts) |
| 396 | if not anns: | 428 | if not anns: |
| 397 | return None | 429 | return None |
| @@ -422,7 +454,7 @@ def _pure_annotation_change_exempt( | |||
| 422 | """ | 454 | """ |
| 423 | if ann_ctx is None: | 455 | if ann_ctx is None: |
| 424 | return False | 456 | return False |
| 425 | - scopes, class_map, has_future = ann_ctx | 457 | + scopes, class_map, has_future, _known = ann_ctx |
| 426 | del_anns = _parse_pure_annotations([t for _, t in del_lines]) | 458 | del_anns = _parse_pure_annotations([t for _, t in del_lines]) |
| 427 | if not del_anns: | 459 | if not del_anns: |
| 428 | return False | 460 | return False |
| @@ -446,6 +478,161 @@ def _pure_annotation_change_exempt( | |||
| 446 | return not _class_consumes_annotations(node, class_map) | 478 | return not _class_consumes_annotations(node, class_map) |
| 447 | 479 | ||
| 448 | 480 | ||
| 481 | +def _parse_assignments(texts) -> list[tuple[ast.stmt, list[str]]] | None: | ||
| 482 | + """Parse diff lines as a block of value-carrying simple-name assignments. | ||
| 483 | + | ||
| 484 | + Returns [(stmt, target_names), ...] when every meaningful (non-blank, | ||
| 485 | + non-comment) line forms assignments (``x = v`` or ``x: T = v``) whose | ||
| 486 | + targets are all plain names; None when the block is unparsable or | ||
| 487 | + contains anything else (bare annotations, attribute/subscript targets, | ||
| 488 | + tuple unpacking, any other statement). | ||
| 489 | + """ | ||
| 490 | + lines = [t for t in texts if t.strip() and not t.strip().startswith("#")] | ||
| 491 | + if not lines: | ||
| 492 | + return None | ||
| 493 | + block = textwrap.dedent("\n".join(lines)) | ||
| 494 | + try: | ||
| 495 | + tree = ast.parse(block) | ||
| 496 | + except (SyntaxError, ValueError): | ||
| 497 | + return None | ||
| 498 | + stmts = [] | ||
| 499 | + for stmt in tree.body: | ||
| 500 | + if isinstance(stmt, ast.Assign): | ||
| 501 | + if not all(isinstance(t, ast.Name) for t in stmt.targets): | ||
| 502 | + return None | ||
| 503 | + stmts.append((stmt, [t.id for t in stmt.targets])) | ||
| 504 | + elif isinstance(stmt, ast.AnnAssign): | ||
| 505 | + if stmt.value is None or not isinstance(stmt.target, ast.Name): | ||
| 506 | + return None | ||
| 507 | + stmts.append((stmt, [stmt.target.id])) | ||
| 508 | + else: | ||
| 509 | + return None | ||
| 510 | + return stmts | ||
| 511 | + | ||
| 512 | + | ||
| 513 | +_LAZY_RHS_NODES = ( | ||
| 514 | + ast.Constant, | ||
| 515 | + ast.Name, | ||
| 516 | + ast.Attribute, | ||
| 517 | + ast.Tuple, | ||
| 518 | + ast.List, | ||
| 519 | + ast.Set, | ||
| 520 | + ast.Dict, | ||
| 521 | + ast.BinOp, | ||
| 522 | + ast.UnaryOp, | ||
| 523 | + ast.BoolOp, | ||
| 524 | + ast.operator, | ||
| 525 | + ast.unaryop, | ||
| 526 | + ast.boolop, | ||
| 527 | + ast.expr_context, | ||
| 528 | +) | ||
| 529 | + | ||
| 530 | + | ||
| 531 | +def _assign_rhs_lazy(value: ast.expr) -> bool: | ||
| 532 | + """True when the right-hand side is free of runtime side effects: | ||
| 533 | + literals, names, attributes and containers/arithmetic over those only | ||
| 534 | + (no calls, subscripts, comprehensions, lambdas, walrus...). Attribute | ||
| 535 | + access and arithmetic can still trigger property/__add__ hooks - an | ||
| 536 | + accepted risk; dict unpacking ({**d}) is not: it iterates the source | ||
| 537 | + mapping at runtime.""" | ||
| 538 | + for node in ast.walk(value): | ||
| 539 | + if not isinstance(node, _LAZY_RHS_NODES): | ||
| 540 | + return False | ||
| 541 | + if isinstance(node, ast.Dict) and any(k is None for k in node.keys): | ||
| 542 | + return False | ||
| 543 | + return True | ||
| 544 | + | ||
| 545 | + | ||
| 546 | +def _decorator_tail(dec: ast.expr) -> str | None: | ||
| 547 | + """Simple name of a decorator (call decorators use their callee): | ||
| 548 | + 'dataclass' for @dataclass / @dataclasses.dataclass / @dataclass(...).""" | ||
| 549 | + expr = dec.func if isinstance(dec, ast.Call) else dec | ||
| 550 | + if isinstance(expr, ast.Name): | ||
| 551 | + return expr.id | ||
| 552 | + if isinstance(expr, ast.Attribute): | ||
| 553 | + return expr.attr | ||
| 554 | + return None | ||
| 555 | + | ||
| 556 | + | ||
| 557 | +def _class_dataclass_like(cls_node: ast.ClassDef, class_map: dict) -> bool: | ||
| 558 | + """True when the class or a same-file ancestor is decorated with | ||
| 559 | + | ||
| 560 | + whose order matters).""" | ||
| 561 | + stack, seen = [cls_node], set() | ||
| 562 | + while stack: | ||
| 563 | + node = stack.pop() | ||
| 564 | + if id(node) in seen: | ||
| 565 | + continue | ||
| 566 | + seen.add(id(node)) | ||
| 567 | + if any(_decorator_tail(d) == "dataclass" for d in node.decorator_list): | ||
| 568 | + return True | ||
| 569 | + for base in node.bases: | ||
| 570 | + name = _base_simple_name(base) | ||
| 571 | + parent = class_map.get(name) if name else None | ||
| 572 | + if parent is not None and id(parent) not in seen: | ||
| 573 | + stack.append(parent) | ||
| 574 | + return False | ||
| 575 | + | ||
| 576 | + | ||
| 577 | +def _last_field_line(cls_node: ast.ClassDef) -> int | None: | ||
| 578 | + """Last base line of the class's annotated fields (dataclass __init__ | ||
| 579 | + parameter order follows field order); None when the class has none.""" | ||
| 580 | + last = None | ||
| 581 | + for stmt in cls_node.body: | ||
| 582 | + if isinstance(stmt, ast.AnnAssign): | ||
| 583 | + end = stmt.end_lineno or stmt.lineno | ||
| 584 | + last = end if last is None else max(last, end) | ||
| 585 | + return last | ||
| 586 | + | ||
| 587 | + | ||
| 588 | +def _new_binding_insert_reason( | ||
| 589 | + add_texts, | ||
| 590 | + a: int, | ||
| 591 | + b: int, | ||
| 592 | + add_indent: int | None, | ||
| 593 | + ann_ctx: tuple | None, | ||
| 594 | +) -> str | None: | ||
| 595 | + """Skip reason when a pure insertion consists solely of assignments that | ||
| 596 | + bind provably-new names with inert right-hand sides; None to record. | ||
| 597 | + | ||
| 598 | + Base code cannot read a name that does not exist yet, so every reader of | ||
| 599 | + a first-time binding is itself an added line carrying its own anchor - | ||
| 600 | + the anchor here is noise (an import-time fingerprint in module/class | ||
| 601 | + scopes). Gates: no meaningful line ends in a comma (such lines are call | ||
| 602 | + or signature arguments, not statements); the block is all simple-name | ||
| 603 | + assignments; no target is a | ||
| 604 | + dunder or already occurs anywhere in the base file (Name occurrences, | ||
| 605 | + def/class/import/parameter bindings); right-hand sides are inert; and in | ||
| 606 | + class scope the class must be plain or dataclass-like, where plain | ||
| 607 | + assigns are never fields while annotated fields are only exempt when | ||
| 608 | + inserted at or after the last existing field (a middle field insertion | ||
| 609 | + would silently rebind positional constructor calls). | ||
| 610 | + """ | ||
| 611 | + if ann_ctx is None or not add_texts: | ||
| 612 | + return None | ||
| 613 | + effective = [t for t in add_texts if t.strip() and not t.strip().startswith("#")] | ||
| 614 | + if not effective or effective[-1].rstrip().endswith(","): | ||
| 615 | + return None | ||
| 616 | + scopes, class_map, has_future, known_names = ann_ctx | ||
| 617 | + stmts = _parse_assignments(add_texts) | ||
| 618 | + if not stmts: | ||
| 619 | + return None | ||
| 620 | + for stmt, names in stmts: | ||
| 621 | + if any(name.startswith("__") or name in known_names for name in names): | ||
| 622 | + return None | ||
| 623 | + if not _assign_rhs_lazy(stmt.value): | ||
| 624 | + return None | ||
| 625 | + scope = _insertion_scope(scopes, a, b, add_indent) | ||
| 626 | + if isinstance(scope, ast.ClassDef) and _class_consumes_annotations(scope, class_map): | ||
| 627 | + if not _class_dataclass_like(scope, class_map): | ||
| 628 | + return None | ||
| 629 | + if any(isinstance(stmt, ast.AnnAssign) for stmt, _ in stmts): | ||
| 630 | + last_field = _last_field_line(scope) | ||
| 631 | + if last_field is not None and a < last_field: | ||
| 632 | + return None | ||
| 633 | + return "first-time binding of new name(s)" | ||
| 634 | + | ||
| 635 | + | ||
| 449 | def _classify_candidate_pairs( | 636 | def _classify_candidate_pairs( |
| 450 | affected: set[int], | 637 | affected: set[int], |
| 451 | pairs: list[tuple], | 638 | pairs: list[tuple], |
| @@ -460,10 +647,17 @@ def _classify_candidate_pairs( | |||
| 460 | comment/docstring line in the base file AND the added lines are comments or | 647 | comment/docstring line in the base file AND the added lines are comments or |
| 461 | doc prose (not parseable Python), i.e. a pure comment/docstring change. | 648 | doc prose (not parseable Python), i.e. a pure comment/docstring change. |
| 462 | Pure type annotation changes are dropped in every direction (insertion, | 649 | Pure type annotation changes are dropped in every direction (insertion, |
| 463 | - replacement, deletion) when provably inert: function-local annotations are | 650 | + replacement, deletion) when provably inert: function-local annotations |
| 464 | - never evaluated; class-level ones need a plain class (no decorators, no | 651 | + are never evaluated; class-level ones need a plain class (no decorators, |
| 465 | - TypedDict/Protocol/Enum/BaseModel/NamedTuple base) and side-effect-free | 652 | + no TypedDict/Protocol/Enum/BaseModel/NamedTuple base) and side-effect-free |
| 466 | annotation expressions unless the file has future annotations. | 653 | annotation expressions unless the file has future annotations. |
| 654 | + First-time bindings are dropped when a pure insertion (no meaningful | ||
| 655 | + line ends in a comma - those are call/signature arguments) consists | ||
| 656 | + solely of assignments binding simple names that occur nowhere in the | ||
| 657 | + base file with inert right-hand sides (base code cannot read a name that | ||
| 658 | + does not exist yet; readers added in the same PR carry their own anchors); | ||
| 659 | + in dataclass-like classes annotated fields are only dropped when inserted | ||
| 660 | + at or after the last existing field. | ||
| 467 | Candidate pairs: without base content (or non-parseable Python) both sides | 661 | Candidate pairs: without base content (or non-parseable Python) both sides |
| 468 | of each pair are counted, bounded by the hunk. | 662 | of each pair are counted, bounded by the hunk. |
| 469 | """ | 663 | """ |
| @@ -532,6 +726,8 @@ def _classify_candidate_pairs( | |||
| 532 | reason = "belongs to a newly added function/class" | 726 | reason = "belongs to a newly added function/class" |
| 533 | if reason is None: | 727 | if reason is None: |
| 534 | reason = _pure_annotation_insert_reason(add_texts, a, b, add_indent, ann_ctx) | 728 | reason = _pure_annotation_insert_reason(add_texts, a, b, add_indent, ann_ctx) |
| 729 | + if reason is None: | ||
| 730 | + reason = _new_binding_insert_reason(add_texts, a, b, add_indent, ann_ctx) | ||
| 535 | # kind == 'blank': isolated blank deletion == one-line insertion | 731 | # kind == 'blank': isolated blank deletion == one-line insertion |
| 536 | elif _between_definitions(a, b, end_lines, start_lines, blanks): | 732 | elif _between_definitions(a, b, end_lines, start_lines, blanks): |
| 537 | reason = "between function/class definitions" | 733 | reason = "between function/class definitions" |
| @@ -11,7 +11,7 @@ __all__ = ["RepoAdapter", "VllmAscendAdapter", "SglangAdapter", "TorchNpuAdapter | |||
| 11 | _ADAPTERS = { | 11 | _ADAPTERS = { |
| 12 | "vllm_ascend": VllmAscendAdapter, | 12 | "vllm_ascend": VllmAscendAdapter, |
| 13 | "sglang": SglangAdapter, | 13 | "sglang": SglangAdapter, |
| 14 | - "torch_npu": TorchNpuAdapter, | 14 | + "pytorch": TorchNpuAdapter, |
| 15 | } | 15 | } |
| 16 | 16 | ||
| 17 | AVAILABLE_REPOS = tuple(_ADAPTERS.keys()) | 17 | AVAILABLE_REPOS = tuple(_ADAPTERS.keys()) |
| @@ -21,7 +21,7 @@ def get_adapter(repo_name: str) -> RepoAdapter: | |||
| 21 | """按仓库名获取适配器实例。 | 21 | """按仓库名获取适配器实例。 |
| 22 | 22 | ||
| 23 | Args: | 23 | Args: |
| 24 | - repo_name: 仓库名(vllm_ascend / sglang / torch_npu) | 24 | + repo_name: 仓库名(vllm_ascend / sglang / pytorch) |
| 25 | 25 | ||
| 26 | Returns: | 26 | Returns: |
| 27 | RepoAdapter 实例 | 27 | RepoAdapter 实例 |
| @@ -21,7 +21,7 @@ class RepoAdapter(ABC): | |||
| 21 | #: 覆盖率数据目录下的测试用例文件夹命名前缀;None 表示无固定前缀(vllm 用 tests__ 前缀 + cpu-ut) | 21 | #: 覆盖率数据目录下的测试用例文件夹命名前缀;None 表示无固定前缀(vllm 用 tests__ 前缀 + cpu-ut) |
| 22 | test_case_dir_prefix: str | None = None | 22 | test_case_dir_prefix: str | None = None |
| 23 | #: 覆盖率文件名 glob 模式(测试目录或其 covdata/ 子目录下)。 | 23 | #: 覆盖率文件名 glob 模式(测试目录或其 covdata/ 子目录下)。 |
| 24 | - #: 默认 "coverage*" 兼容 pytest-cov 裸文件(torch_npu: coverage)与原始格式 | 24 | + #: 默认 "coverage*" 兼容 pytest-cov 裸文件(pytorch: coverage)与原始格式 |
| 25 | #: 带后缀文件(vllm/sglang: coverage.linux-...-workflow.xxx)。 | 25 | #: 带后缀文件(vllm/sglang: coverage.linux-...-workflow.xxx)。 |
| 26 | coverage_file_glob: str = "coverage*" | 26 | coverage_file_glob: str = "coverage*" |
| 27 | 27 | ||
| @@ -1,14 +1,16 @@ | |||
| 1 | -"""torch_npu 仓库适配器(PyTorch / GitCode 社区)。 | 1 | +"""torch_npu 仓库适配器(检测上游 PyTorch:https://github.com/pytorch/pytorch)。 |
| 2 | 2 | ||
| 3 | -- REPO_NAME = "torch_npu",产品代码前缀 "torch_npu/" | 3 | +- REPO_NAME = "torch",产品代码前缀 "torch/",PR 源为 GitHub(--github-pr pytorch/pytorch#N) |
| 4 | -- 测试用例目录:目录名含 "__" 或以 "test_" 开头(如 test__xxx 编码路径), | 4 | +- 测试用例目录:目录名含 "__" 或以 "test_" 开头(如 _inductor__test_add), |
| 5 | 覆盖率文件位于 covdata/ 子目录 | 5 | 覆盖率文件位于 covdata/ 子目录 |
| 6 | - 测试文件规则:test/ 下 test_*.py | 6 | - 测试文件规则:test/ 下 test_*.py |
| 7 | - 测试名规范化:-- -> ::,__ -> /,文件级补 .py | 7 | - 测试名规范化:-- -> ::,__ -> /,文件级补 .py |
| 8 | - 粒度默认值跟随合并版:line=True, function=True | 8 | - 粒度默认值跟随合并版:line=True, function=True |
| 9 | -- 源文件解析使用基类默认候选(source_dir/torch_npu/<rel>,运行需 -s 指向 covstub) | 9 | +- 源文件解析使用基类默认候选(source_dir/torch/<rel>,运行需 -s 指向 covstub, |
| 10 | + covstub/torch/ 对应 pytorch/pytorch 仓库根目录) | ||
| 10 | """ | 11 | """ |
| 11 | 12 | ||
| 13 | +import re | ||
| 12 | from fnmatch import fnmatchcase | 14 | from fnmatch import fnmatchcase |
| 13 | from pathlib import PurePosixPath | 15 | from pathlib import PurePosixPath |
| 14 | 16 | ||
| @@ -18,10 +20,14 @@ from .base import RepoAdapter | |||
| 18 | class TorchNpuAdapter(RepoAdapter): | 20 | class TorchNpuAdapter(RepoAdapter): |
| 19 | """torch_npu 仓库适配器。""" | 21 | """torch_npu 仓库适配器。""" |
| 20 | 22 | ||
| 21 | - repo_name = "torch_npu" | 23 | + repo_name = "torch" |
| 22 | - product_code_prefix = "torch_npu/" | 24 | + product_code_prefix = "torch/" |
| 23 | test_case_dir_prefix = None | 25 | test_case_dir_prefix = None |
| 24 | 26 | ||
| 27 | + #: 精确匹配路径中独立的 torch/ 段(前为路径起点或 /), | ||
| 28 | + #: 避免 pytorch/、torchvision/、test_torch.py 等含 torch 子串的路径被误剥 | ||
| 29 | + _TORCH_PATH_RE = re.compile(r"(?:^|/)torch/(.+)") | ||
| 30 | + | ||
| 25 | #: 可纳入 PR 检测的测试根目录 | 31 | #: 可纳入 PR 检测的测试根目录 |
| 26 | TEST_ROOTS = ("test/",) | 32 | TEST_ROOTS = ("test/",) |
| 27 | 33 | ||
| @@ -30,7 +36,7 @@ class TorchNpuAdapter(RepoAdapter): | |||
| 30 | # ------------------------------------------------------------------ | 36 | # ------------------------------------------------------------------ |
| 31 | def is_test_case_dir(self, name: str) -> bool: | 37 | def is_test_case_dir(self, name: str) -> bool: |
| 32 | # PyTorch: "__" in name or name.startswith("test_") | 38 | # PyTorch: "__" in name or name.startswith("test_") |
| 33 | - return ("__" in name) or name.startswith("test_") | 39 | + return name.endswith(".py") or ("__" in name) or name.startswith("test_") |
| 34 | 40 | ||
| 35 | def normalize_test_name(self, test_name: str) -> str: | 41 | def normalize_test_name(self, test_name: str) -> str: |
| 36 | """ | 42 | """ |
| @@ -40,19 +46,16 @@ class TorchNpuAdapter(RepoAdapter): | |||
| 40 | """ | 46 | """ |
| 41 | # 先转换 -- 为 ::(函数级测试标记) | 47 | # 先转换 -- 为 ::(函数级测试标记) |
| 42 | result = test_name.replace("--", "::") | 48 | result = test_name.replace("--", "::") |
| 43 | - # 再转换 __ 为 / | 49 | + if result.endswith(".py") or (".py::" in result): |
| 44 | - result = result.replace("__", "/") | 50 | + if not result.startswith("test/"): |
| 45 | - # 无 ::(文件级测试)则补 .py 后缀 | 51 | + result = f"test/{result}" |
| 46 | - if "::" not in result: | 52 | + return result |
| 47 | - result = result + ".py" | ||
| 48 | return result | 53 | return result |
| 49 | 54 | ||
| 50 | def covered_path_to_rel(self, path: str) -> str | None: | 55 | def covered_path_to_rel(self, path: str) -> str | None: |
| 51 | - """覆盖率库路径包含 torch_npu/ 时剥离前缀得到相对路径,否则返回 None。""" | 56 | + """覆盖率库路径含独立的 torch/ 段时剥离前缀得到相对路径,否则返回 None。""" |
| 52 | - if self.repo_name not in path: | 57 | + match = self._TORCH_PATH_RE.search(path) |
| 53 | - return None | 58 | + return match.group(1) if match else None |
| 54 | - marker = f"{self.repo_name}/" | ||
| 55 | - return path.split(marker)[-1] if marker in path else path | ||
| 56 | 59 | ||
| 57 | # ------------------------------------------------------------------ | 60 | # ------------------------------------------------------------------ |
| 58 | # 测试文件检测(PR 内容检测) | 61 | # 测试文件检测(PR 内容检测) |
| @@ -8,10 +8,10 @@ | |||
| 8 | 8 | ||
| 9 | ```bash | 9 | ```bash |
| 10 | python -m test_selector --repo vllm_ascend --github-pr "vllm-project/vllm-ascend#12379" | 10 | python -m test_selector --repo vllm_ascend --github-pr "vllm-project/vllm-ascend#12379" |
| 11 | -python -m test_selector --repo torch_npu --gitcode-pr "Ascend/pytorch#46780" | 11 | +python -m test_selector --repo torch_npu --github-pr "pytorch/pytorch#150000" |
| 12 | ``` | 12 | ``` |
| 13 | 13 | ||
| 14 | -> **PR 源选择**:vllm_ascend / sglang 使用 GitHub(`--github-pr`),torch_npu 使用 GitCode(`--gitcode-pr`)。 | 14 | +> **PR 源选择**:三仓库均使用 GitHub(`--github-pr`);torch_npu 检测上游仓库 [pytorch/pytorch](https://github.com/pytorch/pytorch)。 |
| 15 | 15 | ||
| 16 | **核心工作流**(main 函数结构): | 16 | **核心工作流**(main 函数结构): |
| 17 | 17 | ||
| @@ -27,7 +27,7 @@ python -m test_selector --repo torch_npu --gitcode-pr "Ascend/pytorch#46780" | |||
| 27 | ├─────────────────────────────────────────────────────────────────┤ | 27 | ├─────────────────────────────────────────────────────────────────┤ |
| 28 | │ 动作2: 全量触发变更检测(adapter.has_full_suite_changes) │ | 28 | │ 动作2: 全量触发变更检测(adapter.has_full_suite_changes) │ |
| 29 | │ - vllm_ascend: csrc/ 目录变更(非 .md)→ 全量测试 │ | 29 | │ - vllm_ascend: csrc/ 目录变更(非 .md)→ 全量测试 │ |
| 30 | -│ - sglang / torch_npu: 恒为 False(原生/子模块变更走精准匹配) │ | 30 | +│ - sglang / torch_npu: 恒为 False(原生代码变更不解析进精准匹配) │ |
| 31 | ├─────────────────────────────────────────────────────────────────┤ | 31 | ├─────────────────────────────────────────────────────────────────┤ |
| 32 | │ 动作3: 产品代码变更检测(change_detector) │ | 32 | │ 动作3: 产品代码变更检测(change_detector) │ |
| 33 | │ - parse_pr_diff_file() → changed_files_with_lines, renames │ | 33 | │ - parse_pr_diff_file() → changed_files_with_lines, renames │ |
| @@ -49,7 +49,7 @@ python -m test_selector --repo torch_npu --gitcode-pr "Ascend/pytorch#46780" | |||
| 49 | - **精准匹配**:无全量触发时,基于行/函数/文件级匹配选择测试 | 49 | - **精准匹配**:无全量触发时,基于行/函数/文件级匹配选择测试 |
| 50 | - **文件重命名检测**:rename 新路径在覆盖率数据中不存在,用旧路径做文件级匹配召回 | 50 | - **文件重命名检测**:rename 新路径在覆盖率数据中不存在,用旧路径做文件级匹配召回 |
| 51 | - **新增/删除处理**:新增测试文件直接加入;已删除测试文件从结果中移除 | 51 | - **新增/删除处理**:新增测试文件直接加入;已删除测试文件从结果中移除 |
| 52 | -- **排除变更场景**:纯注释/docstring、纯类型注解、函数/类定义间空行插入、新增 def/class 整体、新文件等不构成代码变更的场景,在 diff 解析阶段剔除(见第五章) | 52 | +- **排除变更场景**:纯注释/docstring、纯类型注解、首次新增的变量绑定、函数/类定义间空行插入、新增 def/class 整体、新文件等不构成代码变更的场景,在 diff 解析阶段剔除(见第五章) |
| 53 | 53 | ||
| 54 | **代码结构**: | 54 | **代码结构**: |
| 55 | 55 | ||
| @@ -59,7 +59,7 @@ test_selector/ | |||
| 59 | ├── __main__.py # python -m test_selector 入口 | 59 | ├── __main__.py # python -m test_selector 入口 |
| 60 | ├── github.py # PR 拉取(公共) | 60 | ├── github.py # PR 拉取(公共) |
| 61 | ├── pr_detector.py # 检测 PR 内容:测试文件/删除文件/全量触发(规则委托 adapter) | 61 | ├── pr_detector.py # 检测 PR 内容:测试文件/删除文件/全量触发(规则委托 adapter) |
| 62 | -├── diff_parser.py # 排除变更场景:纯注释/docstring/纯注解/新文件跳过等 | 62 | +├── diff_parser.py # 排除变更场景:纯注释/docstring/纯注解/首次绑定/新文件跳过等 |
| 63 | ├── noise_filter.py # map 噪音处理:import/def/class/docstring/空行 | 63 | ├── noise_filter.py # map 噪音处理:import/def/class/docstring/空行 |
| 64 | ├── function_parser.py # 函数区间解析(sglang 优化版,区间线性扫描) | 64 | ├── function_parser.py # 函数区间解析(sglang 优化版,区间线性扫描) |
| 65 | ├── coverage_selector.py # 构建 测试用例→覆盖文件/行号 映射 | 65 | ├── coverage_selector.py # 构建 测试用例→覆盖文件/行号 映射 |
| @@ -69,7 +69,7 @@ test_selector/ | |||
| 69 | ├── base.py # RepoAdapter 抽象接口 | 69 | ├── base.py # RepoAdapter 抽象接口 |
| 70 | ├── vllm_ascend.py # vllm_ascend 适配器 | 70 | ├── vllm_ascend.py # vllm_ascend 适配器 |
| 71 | ├── sglang.py # sglang 适配器 | 71 | ├── sglang.py # sglang 适配器 |
| 72 | - ├── torch_npu.py # torch_npu(PyTorch / GitCode)适配器 | 72 | + ├── torch_npu.py # torch_npu(上游 PyTorch / GitHub)适配器 |
| 73 | └── __init__.py # 适配器注册表 get_adapter() | 73 | └── __init__.py # 适配器注册表 get_adapter() |
| 74 | ``` | 74 | ``` |
| 75 | 75 | ||
| @@ -97,22 +97,22 @@ test_selector/ | |||
| 97 | │ ├── <sglang 布局>/ | 97 | │ ├── <sglang 布局>/ |
| 98 | │ │ └── ____w__sglang__sglang__test__registered__npu__...__test_xxx/ | 98 | │ │ └── ____w__sglang__sglang__test__registered__npu__...__test_xxx/ |
| 99 | │ │ └── coverage.linux-* # sglang:覆盖率文件直接放在测试目录下 | 99 | │ │ └── coverage.linux-* # sglang:覆盖率文件直接放在测试目录下 |
| 100 | -│ └── <torch_npu 布局>/ | 100 | +│ └── <pytorch 布局>/ |
| 101 | -│ └── _inductor__test_add/ # torch_npu:目录名含 __ 或以 test_ 开头 | 101 | +│ └── _inductor__test_add/ # pytorch:目录名含 __ 或以 test_ 开头 |
| 102 | -│ └── covdata/ # torch_npu:覆盖率文件在 covdata/ 子目录 | 102 | +│ └── covdata/ # pytorch:覆盖率文件在 covdata/ 子目录 |
| 103 | │ └── coverage.* | 103 | │ └── coverage.* |
| 104 | │ | 104 | │ |
| 105 | ├── covstub/ # 源码目录(函数级匹配需要) | 105 | ├── covstub/ # 源码目录(函数级匹配需要) |
| 106 | │ ├── vllm_ascend/ # vllm:对应仓库根目录 | 106 | │ ├── vllm_ascend/ # vllm:对应仓库根目录 |
| 107 | │ ├── sglang/ # sglang:对应仓库 python/sglang/ | 107 | │ ├── sglang/ # sglang:对应仓库 python/sglang/ |
| 108 | -│ └── torch_npu/ # torch_npu:对应 PyTorch 仓库根目录 | 108 | +│ └── torch/ # torch_npu:对应 pytorch/pytorch 仓库根目录 |
| 109 | ├── test_case_map.json # 自动生成的映射文件 | 109 | ├── test_case_map.json # 自动生成的映射文件 |
| 110 | └── recommended_pytest_paths.txt # 推荐测试用例列表(输出) | 110 | └── recommended_pytest_paths.txt # 推荐测试用例列表(输出) |
| 111 | ``` | 111 | ``` |
| 112 | 112 | ||
| 113 | **覆盖率目录布局自动探测**(三仓库兼容): | 113 | **覆盖率目录布局自动探测**(三仓库兼容): |
| 114 | 114 | ||
| 115 | -- 测试目录下存在 `covdata/` 子目录(vllm / torch_npu 布局)→ 从 `covdata/` 读取 `coverage.*` | 115 | +- 测试目录下存在 `covdata/` 子目录(vllm / pytorch 布局)→ 从 `covdata/` 读取 `coverage.*` |
| 116 | - 无 `covdata/` 子目录(sglang 布局)→ 直接读取测试目录下的 `coverage.*` | 116 | - 无 `covdata/` 子目录(sglang 布局)→ 直接读取测试目录下的 `coverage.*` |
| 117 | 117 | ||
| 118 | **测试用例目录识别规则**(adapter 提供): | 118 | **测试用例目录识别规则**(adapter 提供): |
| @@ -137,7 +137,7 @@ test_selector/ | |||
| 137 | | ----------- | ---------------- | ----------------------------------------------------------------- | | 137 | | ----------- | ---------------- | ----------------------------------------------------------------- | |
| 138 | | vllm_ascend | `vllm_ascend/` | `vllm_ascend/core/worker.py` → `core/worker.py` | | 138 | | vllm_ascend | `vllm_ascend/` | `vllm_ascend/core/worker.py` → `core/worker.py` | |
| 139 | | sglang | `python/sglang/` | `python/sglang/srt/models/qwen3_vl.py` → `srt/models/qwen3_vl.py` | | 139 | | sglang | `python/sglang/` | `python/sglang/srt/models/qwen3_vl.py` → `srt/models/qwen3_vl.py` | |
| 140 | -| torch_npu | `torch_npu/` | `torch_npu/contrib/xxx.py` → `contrib/xxx.py` | | 140 | +| torch_npu | `torch/` | `torch/distributed/utils.py` → `distributed/utils.py` | |
| 141 | 141 | ||
| 142 | --- | 142 | --- |
| 143 | 143 | ||
| @@ -189,9 +189,9 @@ python -m test_selector --repo sglang \ | |||
| 189 | --github-pr "sgl-project/sglang#37043" \ | 189 | --github-pr "sgl-project/sglang#37043" \ |
| 190 | --source-dir ./covstub | 190 | --source-dir ./covstub |
| 191 | 191 | ||
| 192 | -# torch_npu(GitCode) | 192 | +# torch_npu(GitHub 上游 pytorch/pytorch) |
| 193 | python -m test_selector --repo torch_npu \ | 193 | python -m test_selector --repo torch_npu \ |
| 194 | - --gitcode-pr "Ascend/pytorch#46780" \ | 194 | + --github-pr "pytorch/pytorch#150000" \ |
| 195 | --source-dir ./covstub | 195 | --source-dir ./covstub |
| 196 | ``` | 196 | ``` |
| 197 | 197 | ||
| @@ -205,18 +205,18 @@ python -m test_selector --repo torch_npu \ | |||
| 205 | 205 | ||
| 206 | ### 公共参数 | 206 | ### 公共参数 |
| 207 | 207 | ||
| 208 | -| 参数 | 必填 | 说明 | | 208 | +| 参数 | 必填 | 说明 | |
| 209 | -| ----------------------- | ----------- | ------------------------------------------------------------------------------------------------------------ | | 209 | +| ----------------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------- | |
| 210 | -| `--repo` / `-r` | 否 | 仓库适配器:`vllm_ascend` / `sglang` / `torch_npu`(默认:`vllm_ascend`,薄入口默认各自仓库) | | 210 | +| `--repo` / `-r` | 否 | 仓库适配器:`vllm_ascend` / `sglang` / `torch_npu`(默认:`vllm_ascend`,薄入口默认各自仓库) | |
| 211 | -| `--github-pr` / `-pr` | 二选一 | GitHub PR(vllm_ascend / sglang),格式:`owner/repo#pr_number` 或仅 `pr_number`(自动从 git remote 推导) | | 211 | +| `--github-pr` / `-pr` | 二选一 | GitHub PR(三仓库通用),格式:`owner/repo#pr_number` 或仅 `pr_number`(自动从 git remote 推导);torch_npu 传上游 `pytorch/pytorch#N` | |
| 212 | -| `--gitcode-pr` | 二选一 | GitCode PR(torch_npu),格式:`owner/repo#pr_number` 或仅 `pr_number`(自动从 git remote 推导);两参数互斥 | | 212 | +| `--gitcode-pr` | 二选一 | GitCode PR(公共备用模块,当前无仓库默认使用),格式:`owner/repo#pr_number` 或仅 `pr_number`(自动从 git remote 推导);两参数互斥 | |
| 213 | -| `--source-dir` / `-s` | 建议 | 源码目录(默认:`covstub`;函数级匹配与噪音过滤需要) | | 213 | +| `--source-dir` / `-s` | 建议 | 源码目录(默认:`covstub`;函数级匹配与噪音过滤需要) | |
| 214 | -| `--map-file` / `-m` | 否 | 映射文件(默认:`test_case_map.json`) | | 214 | +| `--map-file` / `-m` | 否 | 映射文件(默认:`test_case_map.json`) | |
| 215 | -| `--coverage-dir` / `-c` | 构建 map 时 | 覆盖率数据目录(默认:`coverage`) | | 215 | +| `--coverage-dir` / `-c` | 构建 map 时 | 覆盖率数据目录(默认:`coverage`) | |
| 216 | -| `--build-map` / `-b` | 否 | 强制重建测试用例映射;仅构建 map 时使用 | | 216 | +| `--build-map` / `-b` | 否 | 强制重建测试用例映射;仅构建 map 时使用 | |
| 217 | -| `--min-affected` / `-a` | 否 | 最少受影响行数阈值(默认:1) | | 217 | +| `--min-affected` / `-a` | 否 | 最少受影响行数阈值(默认:1) | |
| 218 | -| `--dedup` | 否 | 去重:相同覆盖行的测试只保留一个(默认关闭) | | 218 | +| `--dedup` | 否 | 去重:相同覆盖行的测试只保留一个(默认关闭) | |
| 219 | -| `--skip-imports` | 否 | 函数级匹配时跳过 import 语句行(默认关闭) | | 219 | +| `--skip-imports` | 否 | 函数级匹配时跳过 import 语句行(默认关闭) | |
| 220 | 220 | ||
| 221 | ### 匹配粒度开关 | 221 | ### 匹配粒度开关 |
| 222 | 222 | ||
| @@ -242,12 +242,12 @@ if args.disable_function_match: | |||
| 242 | 242 | ||
| 243 | | 维度 | vllm_ascend | sglang | torch_npu | | 243 | | 维度 | vllm_ascend | sglang | torch_npu | |
| 244 | | -------------- | ----------------------------------------------------- | ----------------------------------------------------------------- | ----------------------------------- | | 244 | | -------------- | ----------------------------------------------------- | ----------------------------------------------------------------- | ----------------------------------- | |
| 245 | -| PR 源 | GitHub | GitHub | GitCode | | 245 | +| PR 源 | GitHub | GitHub | GitHub(pytorch/pytorch) | |
| 246 | -| 产品代码前缀 | `vllm_ascend/` | `python/sglang/` | `torch_npu/` | | 246 | +| 产品代码前缀 | `vllm_ascend/` | `python/sglang/` | `torch/` | |
| 247 | | 测试目录识别 | `tests__` 前缀 / `cpu-ut` | `____w__sglang__sglang__test__` 前缀 | 目录名含 `__` 或以 `test_` 开头 | | 247 | | 测试目录识别 | `tests__` 前缀 / `cpu-ut` | `____w__sglang__sglang__test__` 前缀 | 目录名含 `__` 或以 `test_` 开头 | |
| 248 | | 覆盖率文件位置 | `covdata/` 子目录 | 测试目录下直接放置(兼容探测 covdata) | `covdata/` 子目录 | | 248 | | 覆盖率文件位置 | `covdata/` 子目录 | 测试目录下直接放置(兼容探测 covdata) | `covdata/` 子目录 | |
| 249 | | 测试文件规则 | `tests/e2e/pull_request/`、`tests/ut/` 下 `test_*.py` | `test/registered/`、`test/{unit,e2e,integration}/` 下 `test_*.py` | `test/` 下 `test_*.py` | | 249 | | 测试文件规则 | `tests/e2e/pull_request/`、`tests/ut/` 下 `test_*.py` | `test/registered/`、`test/{unit,e2e,integration}/` 下 `test_*.py` | `test/` 下 `test_*.py` | |
| 250 | -| 全量触发变更 | csrc/ 目录(非 .md)→ 全量测试 | 无(csrc/rust 走精准匹配) | 无(submodule/原生变更走精准匹配) | | 250 | +| 全量触发变更 | csrc/ 目录(非 .md)→ 全量测试 | 无(csrc/rust 走精准匹配) | 无(原生变更走精准匹配) | |
| 251 | | 测试名规范化 | `--`→`::`、`__`→`/`、文件级补 `.py` | 剥离编码前缀恢复 `test/`、`__`→`/`、`--`→`::` | `--`→`::`、`__`→`/`、文件级补 `.py` | | 251 | | 测试名规范化 | `--`→`::`、`__`→`/`、文件级补 `.py` | 剥离编码前缀恢复 `test/`、`__`→`/`、`--`→`::` | `--`→`::`、`__`→`/`、文件级补 `.py` | |
| 252 | | 变更检测 | 本地哈希比对(默认)或 PR diff | 同左 | 同左 | | 252 | | 变更检测 | 本地哈希比对(默认)或 PR diff | 同左 | 同左 | |
| 253 | 253 | ||
| @@ -270,21 +270,22 @@ if args.disable_function_match: | |||
| 270 | 270 | ||
| 271 | > sglang 不触发全量:`csrc/*.cu`、`rust/*.rs` 等原生代码变更不会被解析进精准匹配(仅保留 `.py`),推荐结果由 Python 产品代码匹配 + 新增/删除测试文件决定。 | 271 | > sglang 不触发全量:`csrc/*.cu`、`rust/*.rs` 等原生代码变更不会被解析进精准匹配(仅保留 `.py`),推荐结果由 Python 产品代码匹配 + 新增/删除测试文件决定。 |
| 272 | > | 272 | > |
| 273 | -> torch_npu 不触发全量:submodule 指针更新(如 `third_party/torchair/torchair` 的 commit 变更)与原生代码变更均不会被解析进精准匹配。**注意**:submodule 指针更新意味着子模块内部有真实代码变更,但工具无法穿透到子模块 diff,此时可能返回 0 推荐,存在漏测风险,建议人工评估(详见故障排除表)。 | 273 | +> torch_npu 不触发全量:`torch/csrc/` 等原生代码(C++/CUDA)变更不会被解析进精准匹配(仅保留 `.py`)。**注意**:纯原生代码变更的 PR 可能返回 0 推荐,此时存在漏测风险,建议人工评估是否需跑相关用例(详见故障排除表)。 |
| 274 | 274 | ||
| 275 | ### 排除变更场景(diff_parser) | 275 | ### 排除变更场景(diff_parser) |
| 276 | 276 | ||
| 277 | `_parse_diff_base_lines()` 解析统一 diff 文本为受影响的 base(变更前)行号,并在解析阶段/后续分类阶段排除以下**不构成代码变更**的场景: | 277 | `_parse_diff_base_lines()` 解析统一 diff 文本为受影响的 base(变更前)行号,并在解析阶段/后续分类阶段排除以下**不构成代码变更**的场景: |
| 278 | 278 | ||
| 279 | -| 场景 | 处理 | | 279 | +| 场景 | 处理 | |
| 280 | -| ----------------------------------------- | ------------------------------------------------------------ | | 280 | +| ----------------------------------------- | -------------------------------------------------------------------------------- | |
| 281 | -| 纯注释/docstring 变更 | 删除组全为注释/docstring 且新增为注释/doc 说明 → 剔除 | | 281 | +| 纯注释/docstring 变更 | 删除组全为注释/docstring 且新增为注释/doc 说明 → 剔除 | |
| 282 | -| 纯类型注解变更(插入/替换/删除) | 可证明惰性的注解(函数体内、普通类、模块级,无副作用)→ 豁免 | | 282 | +| 纯类型注解变更(插入/替换/删除) | 可证明惰性的注解(函数体内、普通类、模块级,无副作用)→ 豁免 | |
| 283 | -| 函数/类定义之间的空行插入 | 位于两个 def/class 之间 → 排除 | | 283 | +| 首次新增的变量绑定(纯插入) | 全部为绑定全新名字的带值赋值(名字在 base 中不存在、RHS 惰性、无尾逗号行)→ 豁免 | |
| 284 | -| 新增 def/class 整体 | 插入文本属于新定义的函数/类 → 排除 | | 284 | +| 函数/类定义之间的空行插入 | 位于两个 def/class 之间 → 排除 | |
| 285 | -| 新文件(`--- /dev/null` + `@@ -0,0 ...`) | 无 base 版本,跳过行级解析(避免无意义的 base 内容拉取) | | 285 | +| 新增 def/class 整体 | 插入文本属于新定义的函数/类 → 排除 | |
| 286 | -| 函数体内纯插入 | 记录插入位置上方的行号 | | 286 | +| 新文件(`--- /dev/null` + `@@ -0,0 ...`) | 无 base 版本,跳过行级解析(避免无意义的 base 内容拉取) | |
| 287 | -| 隔离空行删除 | 按单行插入处理为候选对(上、下两行) | | 287 | +| 函数体内纯插入 | 记录插入位置上方的行号 | |
| 288 | +| 隔离空行删除 | 按单行插入处理为候选对(上、下两行) | | ||
| 288 | 289 | ||
| 289 | **候选对分类**(`_classify_candidate_pairs`,需要 base 文件内容):通过 GitHub contents API / GitCode raw 接口单路拉取 base 内容后做 AST 分类;拉取失败重试 3 次后直接退出(不允许降级),仅 AST 解析失败等无法分类的场景回退为 hunk 范围内双侧计数。 | 290 | **候选对分类**(`_classify_candidate_pairs`,需要 base 文件内容):通过 GitHub contents API / GitCode raw 接口单路拉取 base 内容后做 AST 分类;拉取失败重试 3 次后直接退出(不允许降级),仅 AST 解析失败等无法分类的场景回退为 hunk 范围内双侧计数。 |
| 290 | 291 | ||
| @@ -387,9 +388,9 @@ for path, label in file_level_paths: # rename 旧路径 + 删除路径 | |||
| 387 | - **PR 模式**:`parse_pr_diff_file()` 解析 PR diff,产出变更行号 + 重命名映射 + 删除列表 | 388 | - **PR 模式**:`parse_pr_diff_file()` 解析 PR diff,产出变更行号 + 重命名映射 + 删除列表 |
| 388 | - **本地模式**(默认):`detect_changes_by_comparison()` 扫描 `--source-dir` 下全部 `.py`,与 `.file_hashes.json` 基线比对 MD5,变更文件保守返回全部行号(1~9999),首次运行生成基线 | 389 | - **本地模式**(默认):`detect_changes_by_comparison()` 扫描 `--source-dir` 下全部 `.py`,与 `.file_hashes.json` 基线比对 MD5,变更文件保守返回全部行号(1~9999),首次运行生成基线 |
| 389 | 390 | ||
| 390 | -### GitCode PR 拉取(torch_npu) | 391 | +### GitCode PR 拉取(公共备用模块) |
| 391 | 392 | ||
| 392 | -torch_npu 的 PR 拉取通过公共模块 `test_selector.gitcode`(`--gitcode-pr`),与 GitHub 版(`test_selector.github`,`--github-pr`)的差异: | 393 | +GitCode PR 拉取通过公共模块 `test_selector.gitcode`(`--gitcode-pr`,当前无仓库默认使用;torch_npu 已切换到 GitHub 上游仓库 pytorch/pytorch),与 GitHub 版(`test_selector.github`,`--github-pr`)的差异: |
| 393 | 394 | ||
| 394 | | 维度 | github.py | gitcode.py | | 395 | | 维度 | github.py | gitcode.py | |
| 395 | | ----------------- | -------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 396 | | ----------------- | -------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | |
| @@ -454,7 +455,7 @@ torch_npu 的 PR 拉取通过公共模块 `test_selector.gitcode`(`--gitcode-p | |||
| 454 | **注意**: | 455 | **注意**: |
| 455 | 456 | ||
| 456 | - map 的 key 为 normalize 后的测试路径(vllm:`tests/.../test_xxx.py` 或 `...::test_func`、`cpu-ut`;sglang:`test/.../test_xxx.py`;torch_npu:`test/.../test_xxx.py`、`_inductor/test_add.py` 等,文件级补 `.py`) | 457 | - map 的 key 为 normalize 后的测试路径(vllm:`tests/.../test_xxx.py` 或 `...::test_func`、`cpu-ut`;sglang:`test/.../test_xxx.py`;torch_npu:`test/.../test_xxx.py`、`_inductor/test_add.py` 等,文件级补 `.py`) |
| 457 | -- files 的 key 为剥离产品代码前缀后的相对路径(vllm:`core/worker.py`;sglang:`srt/xxx.py`;torch_npu:`contrib/xxx.py`) | 458 | +- files 的 key 为剥离产品代码前缀后的相对路径(vllm:`core/worker.py`;sglang:`srt/xxx.py`;torch_npu:`distributed/xxx.py`) |
| 458 | - 保存时强制使用 LF 换行(`newline="\n"`),避免 Windows CRLF 导致 JSON 体积膨胀,保证跨平台 git 差异对比一致 | 459 | - 保存时强制使用 LF 换行(`newline="\n"`),避免 Windows CRLF 导致 JSON 体积膨胀,保证跨平台 git 差异对比一致 |
| 459 | 460 | ||
| 460 | --- | 461 | --- |
| @@ -531,9 +532,9 @@ python -m test_selector --repo vllm_ascend \ | |||
| 531 | python -m test_selector --repo sglang \ | 532 | python -m test_selector --repo sglang \ |
| 532 | --github-pr "sgl-project/sglang#37043" -s ./covstub | 533 | --github-pr "sgl-project/sglang#37043" -s ./covstub |
| 533 | 534 | ||
| 534 | -# GitCode PR(torch_npu) | 535 | +# GitHub PR(torch_npu,上游 pytorch/pytorch) |
| 535 | python -m test_selector --repo torch_npu \ | 536 | python -m test_selector --repo torch_npu \ |
| 536 | - --gitcode-pr "Ascend/pytorch#46780" -s ./covstub | 537 | + --github-pr "pytorch/pytorch#150000" -s ./covstub |
| 537 | ``` | 538 | ``` |
| 538 | 539 | ||
| 539 | ### 场景2:更新覆盖率数据后重建映射 | 540 | ### 场景2:更新覆盖率数据后重建映射 |
| @@ -592,23 +593,23 @@ selected, reason = ts.select_tests(changed, source_dir="covstub") | |||
| 592 | 593 | ||
| 593 | ## 九、故障排除 | 594 | ## 九、故障排除 |
| 594 | 595 | ||
| 595 | -| 问题 | 可能原因 | 解决方案 | | 596 | +| 问题 | 可能原因 | 解决方案 | |
| 596 | -| ------------------------------------------------------------------- | --------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | 597 | +| ------------------------------------------------------------------- | ------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | |
| 597 | -| `Error: --coverage-dir is required when building the test case map` | map 文件不存在且未传覆盖率目录 | 首次使用需传 `--coverage-dir`(或先 `--build-map`) | | 598 | +| `Error: --coverage-dir is required when building the test case map` | map 文件不存在且未传覆盖率目录 | 首次使用需传 `--coverage-dir`(或先 `--build-map`) | |
| 598 | -| `Coverage data directory not found` | 覆盖率目录路径不对 | 确认 `--coverage-dir` 指向含测试用例目录的根目录 | | 599 | +| `Coverage data directory not found` | 覆盖率目录路径不对 | 确认 `--coverage-dir` 指向含测试用例目录的根目录 | |
| 599 | -| `没有测试用例覆盖变更的代码行` | 变更文件无覆盖数据 | 确认覆盖率数据来自变更前的全量用例 | | 600 | +| `没有测试用例覆盖变更的代码行` | 变更文件无覆盖数据 | 确认覆盖率数据来自变更前的全量用例 | |
| 600 | -| `解析 PR 失败` / GitHub API 超时 | 内网无法直连 GitHub | 设置代理环境变量后重试(见下方) | | 601 | +| `解析 PR 失败` / GitHub API 超时 | 内网无法直连 GitHub | 设置代理环境变量后重试(见下方) | |
| 601 | -| `WinError 10060` 连接超时(GitCode) | 内网无法直连 GitCode API | 设置代理环境变量后重试(见下方);git 全局已配置 `proxycn2.huawei.com:8080` 认证代理时,PowerShell 中执行 `$env:HTTPS_PROXY=...` 后重试 | | 602 | +| `WinError 10060` 连接超时(GitCode) | 内网无法直连 GitCode API | 设置代理环境变量后重试(见下方);git 全局已配置 `proxycn2.huawei.com:8080` 认证代理时,PowerShell 中执行 `$env:HTTPS_PROXY=...` 后重试 | |
| 602 | -| `函数级匹配失败` | 源码目录路径不对 | 确认 `--source-dir` 指向包含 `vllm_ascend/` / `sglang/` / `torch_npu/` 包的根目录 | | 603 | +| `函数级匹配失败` | 源码目录路径不对 | 确认 `--source-dir` 指向包含 `vllm_ascend/` / `sglang/` / `torch/` 包的根目录 | |
| 603 | -| 推荐结果过多 | 变更函数被大量测试覆盖 | 启用 `--dedup` 或提高 `--min-affected` | | 604 | +| 推荐结果过多 | 变更函数被大量测试覆盖 | 启用 `--dedup` 或提高 `--min-affected` | |
| 604 | -| 推荐结果为空 | PR 仅含原生代码变更且无新测试文件(sglang) | 正常,原生代码变更不触发全量 | | 605 | +| 推荐结果为空 | PR 仅含原生代码变更且无新测试文件(sglang) | 正常,原生代码变更不触发全量 | |
| 605 | -| 推荐结果为空 | PR 仅含 submodule 指针更新(torch_npu,如 `third_party/torchair/torchair`) | 正常但**存在漏测风险**:submodule commit 变更意味着子模块内部有真实代码变更,工具无法穿透;建议人工评估是否需跑 torchchair 相关用例 | | 606 | +| 推荐结果为空 | PR 仅含原生代码变更(torch_npu,如 `torch/csrc/` 下 C++/CUDA 文件) | 正常但**存在漏测风险**:原生代码变更不解析进精准匹配;建议人工评估是否需跑相关用例 | |
| 606 | -| `Full-Suite Changes Detected`(vllm) | diff 含 csrc/ 目录变更 | 正常行为,触发全量测试 | | 607 | +| `Full-Suite Changes Detected`(vllm) | diff 含 csrc/ 目录变更 | 正常行为,触发全量测试 | |
| 607 | -| `New Test Files Added: X` | diff 中有新增测试文件 | 正常,直接加入推荐列表 | | 608 | +| `New Test Files Added: X` | diff 中有新增测试文件 | 正常,直接加入推荐列表 | |
| 608 | -| `Deleted Test Files Removed: X` | diff 中有已删除测试文件 | 正常,从推荐列表移除 | | 609 | +| `Deleted Test Files Removed: X` | diff 中有已删除测试文件 | 正常,从推荐列表移除 | |
| 609 | -| `Detected X Renamed File(s)` | diff 中包含文件重命名 | 正常,旧路径做文件级匹配 | | 610 | +| `Detected X Renamed File(s)` | diff 中包含文件重命名 | 正常,旧路径做文件级匹配 | |
| 610 | -| 变更行数偏多、含 context 行 | diff 解析按 hunk 范围推算 | 已知限制,算法保守偏多报 | | 611 | +| 变更行数偏多、含 context 行 | diff 解析按 hunk 范围推算 | 已知限制,算法保守偏多报 | |
| 611 | -| `Unknown repo 'xxx'` | `--repo` 拼写错误 | 使用 `vllm_ascend` / `sglang` / `torch_npu` | | 612 | +| `Unknown repo 'xxx'` | `--repo` 拼写错误 | 使用 `vllm_ascend` / `sglang` / `torch_npu` | |
| 612 | 613 | ||
| 613 | **代理配置示例(华为内网环境)**: | 614 | **代理配置示例(华为内网环境)**: |
| 614 | 615 | ||
| @@ -618,7 +619,7 @@ set HTTP_PROXY=http://<user>:<pwd>@proxycn2.huawei.com:8080/ | |||
| 618 | set HTTPS_PROXY=http://<user>:<pwd>@proxycn2.huawei.com:8080/ | 619 | set HTTPS_PROXY=http://<user>:<pwd>@proxycn2.huawei.com:8080/ |
| 619 | 620 | ||
| 620 | python -m test_selector --repo vllm_ascend --github-pr "vllm-project/vllm-ascend#16104" -s ./covstub | 621 | python -m test_selector --repo vllm_ascend --github-pr "vllm-project/vllm-ascend#16104" -s ./covstub |
| 621 | -python -m test_selector --repo torch_npu --gitcode-pr "Ascend/pytorch#46780" -s ./covstub | 622 | +python -m test_selector --repo torch_npu --github-pr "pytorch/pytorch#150000" -s ./covstub |
| 622 | ``` | 623 | ``` |
| 623 | 624 | ||
| 624 | > **说明**: | 625 | > **说明**: |
| @@ -1,10 +1,10 @@ | |||
| 1 | -"""test_selector - 覆盖率驱动的精准测试选择器(vllm_ascend / sglang / torch_npu 统一架构)。 | 1 | +"""test_selector - 覆盖率驱动的精准测试选择器(vllm_ascend / sglang / pytorch 统一架构)。 |
| 2 | 2 | ||
| 3 | 公共逻辑集中在 test_selector/ 包(diff 解析、噪音过滤、覆盖率选择、CLI 等), | 3 | 公共逻辑集中在 test_selector/ 包(diff 解析、噪音过滤、覆盖率选择、CLI 等), |
| 4 | 仓库差异通过 test_selector.repos.RepoAdapter 注入,保证架构统一、差异按仓库归类。 | 4 | 仓库差异通过 test_selector.repos.RepoAdapter 注入,保证架构统一、差异按仓库归类。 |
| 5 | 5 | ||
| 6 | 入口: | 6 | 入口: |
| 7 | - python -m test_selector --repo <vllm_ascend|sglang|torch_npu> ... | 7 | + python -m test_selector --repo <vllm_ascend|sglang|pytorch> ... |
| 8 | """ | 8 | """ |
| 9 | 9 | ||
| 10 | from .repos import AVAILABLE_REPOS, RepoAdapter, get_adapter | 10 | from .repos import AVAILABLE_REPOS, RepoAdapter, get_adapter |
| @@ -106,6 +106,13 @@ class CodeChangeDetector: | |||
| 106 | NamedTuple base) or module level, unless the annotation expression | 106 | NamedTuple base) or module level, unless the annotation expression |
| 107 | may have runtime side effects (calls, Annotated metadata) and the | 107 | may have runtime side effects (calls, Annotated metadata) and the |
| 108 | file has no ``from __future__ import annotations``. | 108 | file has no ``from __future__ import annotations``. |
| 109 | + - First-time bindings are excluded for pure insertions (needs base | ||
| 110 | + content): every added line is an assignment binding simple names | ||
| 111 | + that occur nowhere in the base file, with inert right-hand sides; | ||
| 112 | + no meaningful line ends in a comma (such lines are call/signature | ||
| 113 | + arguments, not statements); class-scope insertions follow dataclass | ||
| 114 | + field-order rules (plain assigns are not fields, annotated fields | ||
| 115 | + must land at or after the last existing field). | ||
| 109 | - Isolated blank-line deletion (neighbours not deleted): treated as a | 116 | - Isolated blank-line deletion (neighbours not deleted): treated as a |
| 110 | one-line insertion -> candidate pair (line above, line below). | 117 | one-line insertion -> candidate pair (line above, line below). |
| 111 | - Pure insertions and blank-deletion pairs are classified via ast of | 118 | - Pure insertions and blank-deletion pairs are classified via ast of |
| @@ -10,11 +10,11 @@ | |||
| 10 | 10 | ||
| 11 | 用法: | 11 | 用法: |
| 12 | python -m test_selector --repo vllm_ascend --github-pr owner/repo#123 | 12 | python -m test_selector --repo vllm_ascend --github-pr owner/repo#123 |
| 13 | - python -m test_selector --repo torch_npu --gitcode-pr Ascend/pytorch#123 | 13 | + python -m test_selector --repo torch_npu --github-pr pytorch/pytorch#123 |
| 14 | python -m test_selector --repo sglang --build-map --coverage-dir coverage | 14 | python -m test_selector --repo sglang --build-map --coverage-dir coverage |
| 15 | python vllm/test_selector.py --github-pr 123 # 薄入口(默认 vllm_ascend) | 15 | python vllm/test_selector.py --github-pr 123 # 薄入口(默认 vllm_ascend) |
| 16 | python sglang/test_selector.py --github-pr 123 # 薄入口(默认 sglang) | 16 | python sglang/test_selector.py --github-pr 123 # 薄入口(默认 sglang) |
| 17 | - python PyTorch/test_selector.py --gitcode-pr 123 # 薄入口(默认 torch_npu) | 17 | + python PyTorch/test_selector.py --github-pr 123 # 薄入口(默认 torch_npu,检测上游 pytorch/pytorch) |
| 18 | """ | 18 | """ |
| 19 | 19 | ||
| 20 | import argparse | 20 | import argparse |
| @@ -49,20 +49,21 @@ class CoverageSelector: | |||
| 49 | if not self.coverage_data_dir or not self.coverage_data_dir.exists(): | 49 | if not self.coverage_data_dir or not self.coverage_data_dir.exists(): |
| 50 | print(f" Warning: Coverage data directory not found: {self.coverage_data_dir}") | 50 | print(f" Warning: Coverage data directory not found: {self.coverage_data_dir}") |
| 51 | return test_cases | 51 | return test_cases |
| 52 | - for item in self.coverage_data_dir.iterdir(): | 52 | + for item in self.coverage_data_dir.rglob("*"): |
| 53 | if not item.is_dir(): | 53 | if not item.is_dir(): |
| 54 | continue | 54 | continue |
| 55 | name = item.name | 55 | name = item.name |
| 56 | # 目录名规则由仓库适配器决定(vllm: tests__ 前缀 / cpu-ut;sglang: ____w__... 前缀) | 56 | # 目录名规则由仓库适配器决定(vllm: tests__ 前缀 / cpu-ut;sglang: ____w__... 前缀) |
| 57 | if not self.adapter.is_test_case_dir(name): | 57 | if not self.adapter.is_test_case_dir(name): |
| 58 | continue | 58 | continue |
| 59 | - # 布局探测:covdata/ 子目录(vllm/torch_npu)或目录下直接放置(sglang) | 59 | + # 布局探测:covdata/ 子目录(vllm/pytorch)或目录下直接放置(sglang) |
| 60 | covdata_dir = item / "covdata" | 60 | covdata_dir = item / "covdata" |
| 61 | has_cov_files = any(item.glob(self.adapter.coverage_file_glob)) | 61 | has_cov_files = any(item.glob(self.adapter.coverage_file_glob)) |
| 62 | if covdata_dir.exists() and any(covdata_dir.glob(self.adapter.coverage_file_glob)): | 62 | if covdata_dir.exists() and any(covdata_dir.glob(self.adapter.coverage_file_glob)): |
| 63 | has_cov_files = True | 63 | has_cov_files = True |
| 64 | if has_cov_files: | 64 | if has_cov_files: |
| 65 | - test_cases.append(name) | 65 | + rel = item.relative_to(self.coverage_data_dir).as_posix() |
| 66 | + test_cases.append(rel) | ||
| 66 | return sorted(test_cases) | 67 | return sorted(test_cases) |
| 67 | 68 | ||
| 68 | def normalize_test_name(self, test_name: str) -> str: | 69 | def normalize_test_name(self, test_name: str) -> str: |
| @@ -142,7 +143,7 @@ class CoverageSelector: | |||
| 142 | 143 | ||
| 143 | file_lines_map = defaultdict(set) # filepath -> set of lines | 144 | file_lines_map = defaultdict(set) # filepath -> set of lines |
| 144 | 145 | ||
| 145 | - # 布局探测:vllm/torch_npu 将 coverage 放在 covdata/ 子目录,sglang 直接放在测试目录下 | 146 | + # 布局探测:vllm/pytroch 将 coverage 放在 covdata/ 子目录,sglang 直接放在测试目录下 |
| 146 | cov_dirs = [covdata_dir] if covdata_dir.exists() else [test_case_dir] | 147 | cov_dirs = [covdata_dir] if covdata_dir.exists() else [test_case_dir] |
| 147 | 148 | ||
| 148 | for cov_dir in cov_dirs: | 149 | for cov_dir in cov_dirs: |
| @@ -4,6 +4,8 @@ | |||
| 4 | 代码变更的场景: | 4 | 代码变更的场景: |
| 5 | - 纯注释/docstring 变更 | 5 | - 纯注释/docstring 变更 |
| 6 | - 纯类型注解变更(插入/替换/删除三个方向,可证明惰性时豁免) | 6 | - 纯类型注解变更(插入/替换/删除三个方向,可证明惰性时豁免) |
| 7 | +- 首次新增的变量绑定(纯插入块全部为绑定全新名字的带值赋值,RHS 惰性, | ||
| 8 | + 无尾逗号行——尾逗号行是调用/签名参数而非语句) | ||
| 7 | - 函数/类定义之间的空行插入 | 9 | - 函数/类定义之间的空行插入 |
| 8 | - 新增函数/类整体 | 10 | - 新增函数/类整体 |
| 9 | 11 | ||
| @@ -318,14 +320,27 @@ def _class_consumes_annotations(cls_node: ast.ClassDef, class_map: dict) -> bool | |||
| 318 | return False | 320 | return False |
| 319 | 321 | ||
| 320 | 322 | ||
| 321 | -def _collect_class_info(source: str) -> tuple[list[tuple], dict[str, ast.ClassDef], bool]: | 323 | +def _arg_names(args: ast.arguments) -> set[str]: |
| 322 | - """Parse source and return (scopes, class_map, has_future): | 324 | + """Names bound by a function/lambda parameter list.""" |
| 325 | + names = {a.arg for a in args.posonlyargs + args.args + args.kwonlyargs} | ||
| 326 | + if args.vararg: | ||
| 327 | + names.add(args.vararg.arg) | ||
| 328 | + if args.kwarg: | ||
| 329 | + names.add(args.kwarg.arg) | ||
| 330 | + return names | ||
| 323 | 331 | ||
| 324 | - scopes : [(lineno, end_lineno, col_offset, node)] of every | 332 | + |
| 325 | - function/class definition, for annotation scope resolution; | 333 | +def _collect_class_info(source: str) -> tuple[list[tuple], dict[str, ast.ClassDef], bool, set[str]]: |
| 326 | - class_map : {class name -> ClassDef} for same-file base resolution | 334 | + """Parse source and return (scopes, class_map, has_future, known_names): |
| 327 | - (a later definition shadows an earlier one); | 335 | + |
| 328 | - has_future: the file has ``from __future__ import annotations``. | 336 | + scopes : [(lineno, end_lineno, col_offset, node)] of every |
| 337 | + function/class definition, for annotation scope resolution; | ||
| 338 | + class_map : {class name -> ClassDef} for same-file base resolution | ||
| 339 | + (a later definition shadows an earlier one); | ||
| 340 | + has_future : the file has ``from __future__ import annotations``. | ||
| 341 | + known_names : every name the base file already binds or reads (Name | ||
| 342 | + occurrences, def/class names, import aliases, parameters), | ||
| 343 | + for first-time-binding detection. | ||
| 329 | """ | 344 | """ |
| 330 | tree = ast.parse(source) | 345 | tree = ast.parse(source) |
| 331 | scopes = [ | 346 | scopes = [ |
| @@ -344,7 +359,24 @@ def _collect_class_info(source: str) -> tuple[list[tuple], dict[str, ast.ClassDe | |||
| 344 | and any(alias.name == "annotations" for alias in n.names) | 359 | and any(alias.name == "annotations" for alias in n.names) |
| 345 | for n in tree.body | 360 | for n in tree.body |
| 346 | ) | 361 | ) |
| 347 | - return scopes, class_map, has_future | 362 | + known_names = set() |
| 363 | + for node in ast.walk(tree): | ||
| 364 | + if isinstance(node, ast.Name): | ||
| 365 | + known_names.add(node.id) | ||
| 366 | + elif isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)): | ||
| 367 | + known_names.add(node.name) | ||
| 368 | + known_names.update(_arg_names(node.args)) | ||
| 369 | + elif isinstance(node, ast.ClassDef): | ||
| 370 | + known_names.add(node.name) | ||
| 371 | + elif isinstance(node, ast.Lambda): | ||
| 372 | + known_names.update(_arg_names(node.args)) | ||
| 373 | + elif isinstance(node, ast.ImportFrom): | ||
| 374 | + for alias in node.names: | ||
| 375 | + known_names.add(alias.asname or alias.name) | ||
| 376 | + elif isinstance(node, ast.Import): | ||
| 377 | + for alias in node.names: | ||
| 378 | + known_names.add(alias.asname or alias.name.split(".")[0]) | ||
| 379 | + return scopes, class_map, has_future, known_names | ||
| 348 | 380 | ||
| 349 | 381 | ||
| 350 | def _insertion_scope(scopes: list[tuple], a: int, b: int, add_indent: int | None): | 382 | def _insertion_scope(scopes: list[tuple], a: int, b: int, add_indent: int | None): |
| @@ -391,7 +423,7 @@ def _pure_annotation_insert_reason( | |||
| 391 | """ | 423 | """ |
| 392 | if ann_ctx is None or not add_texts: | 424 | if ann_ctx is None or not add_texts: |
| 393 | return None | 425 | return None |
| 394 | - scopes, class_map, has_future = ann_ctx | 426 | + scopes, class_map, has_future, _known = ann_ctx |
| 395 | anns = _parse_pure_annotations(add_texts) | 427 | anns = _parse_pure_annotations(add_texts) |
| 396 | if not anns: | 428 | if not anns: |
| 397 | return None | 429 | return None |
| @@ -422,7 +454,7 @@ def _pure_annotation_change_exempt( | |||
| 422 | """ | 454 | """ |
| 423 | if ann_ctx is None: | 455 | if ann_ctx is None: |
| 424 | return False | 456 | return False |
| 425 | - scopes, class_map, has_future = ann_ctx | 457 | + scopes, class_map, has_future, _known = ann_ctx |
| 426 | del_anns = _parse_pure_annotations([t for _, t in del_lines]) | 458 | del_anns = _parse_pure_annotations([t for _, t in del_lines]) |
| 427 | if not del_anns: | 459 | if not del_anns: |
| 428 | return False | 460 | return False |
| @@ -446,6 +478,161 @@ def _pure_annotation_change_exempt( | |||
| 446 | return not _class_consumes_annotations(node, class_map) | 478 | return not _class_consumes_annotations(node, class_map) |
| 447 | 479 | ||
| 448 | 480 | ||
| 481 | +def _parse_assignments(texts) -> list[tuple[ast.stmt, list[str]]] | None: | ||
| 482 | + """Parse diff lines as a block of value-carrying simple-name assignments. | ||
| 483 | + | ||
| 484 | + Returns [(stmt, target_names), ...] when every meaningful (non-blank, | ||
| 485 | + non-comment) line forms assignments (``x = v`` or ``x: T = v``) whose | ||
| 486 | + targets are all plain names; None when the block is unparsable or | ||
| 487 | + contains anything else (bare annotations, attribute/subscript targets, | ||
| 488 | + tuple unpacking, any other statement). | ||
| 489 | + """ | ||
| 490 | + lines = [t for t in texts if t.strip() and not t.strip().startswith("#")] | ||
| 491 | + if not lines: | ||
| 492 | + return None | ||
| 493 | + block = textwrap.dedent("\n".join(lines)) | ||
| 494 | + try: | ||
| 495 | + tree = ast.parse(block) | ||
| 496 | + except (SyntaxError, ValueError): | ||
| 497 | + return None | ||
| 498 | + stmts = [] | ||
| 499 | + for stmt in tree.body: | ||
| 500 | + if isinstance(stmt, ast.Assign): | ||
| 501 | + if not all(isinstance(t, ast.Name) for t in stmt.targets): | ||
| 502 | + return None | ||
| 503 | + stmts.append((stmt, [t.id for t in stmt.targets])) | ||
| 504 | + elif isinstance(stmt, ast.AnnAssign): | ||
| 505 | + if stmt.value is None or not isinstance(stmt.target, ast.Name): | ||
| 506 | + return None | ||
| 507 | + stmts.append((stmt, [stmt.target.id])) | ||
| 508 | + else: | ||
| 509 | + return None | ||
| 510 | + return stmts | ||
| 511 | + | ||
| 512 | + | ||
| 513 | +_LAZY_RHS_NODES = ( | ||
| 514 | + ast.Constant, | ||
| 515 | + ast.Name, | ||
| 516 | + ast.Attribute, | ||
| 517 | + ast.Tuple, | ||
| 518 | + ast.List, | ||
| 519 | + ast.Set, | ||
| 520 | + ast.Dict, | ||
| 521 | + ast.BinOp, | ||
| 522 | + ast.UnaryOp, | ||
| 523 | + ast.BoolOp, | ||
| 524 | + ast.operator, | ||
| 525 | + ast.unaryop, | ||
| 526 | + ast.boolop, | ||
| 527 | + ast.expr_context, | ||
| 528 | +) | ||
| 529 | + | ||
| 530 | + | ||
| 531 | +def _assign_rhs_lazy(value: ast.expr) -> bool: | ||
| 532 | + """True when the right-hand side is free of runtime side effects: | ||
| 533 | + literals, names, attributes and containers/arithmetic over those only | ||
| 534 | + (no calls, subscripts, comprehensions, lambdas, walrus...). Attribute | ||
| 535 | + access and arithmetic can still trigger property/__add__ hooks - an | ||
| 536 | + accepted risk; dict unpacking ({**d}) is not: it iterates the source | ||
| 537 | + mapping at runtime.""" | ||
| 538 | + for node in ast.walk(value): | ||
| 539 | + if not isinstance(node, _LAZY_RHS_NODES): | ||
| 540 | + return False | ||
| 541 | + if isinstance(node, ast.Dict) and any(k is None for k in node.keys): | ||
| 542 | + return False | ||
| 543 | + return True | ||
| 544 | + | ||
| 545 | + | ||
| 546 | +def _decorator_tail(dec: ast.expr) -> str | None: | ||
| 547 | + """Simple name of a decorator (call decorators use their callee): | ||
| 548 | + 'dataclass' for @dataclass / @dataclasses.dataclass / @dataclass(...).""" | ||
| 549 | + expr = dec.func if isinstance(dec, ast.Call) else dec | ||
| 550 | + if isinstance(expr, ast.Name): | ||
| 551 | + return expr.id | ||
| 552 | + if isinstance(expr, ast.Attribute): | ||
| 553 | + return expr.attr | ||
| 554 | + return None | ||
| 555 | + | ||
| 556 | + | ||
| 557 | +def _class_dataclass_like(cls_node: ast.ClassDef, class_map: dict) -> bool: | ||
| 558 | + """True when the class or a same-file ancestor is decorated with | ||
| 559 | + | ||
| 560 | + whose order matters).""" | ||
| 561 | + stack, seen = [cls_node], set() | ||
| 562 | + while stack: | ||
| 563 | + node = stack.pop() | ||
| 564 | + if id(node) in seen: | ||
| 565 | + continue | ||
| 566 | + seen.add(id(node)) | ||
| 567 | + if any(_decorator_tail(d) == "dataclass" for d in node.decorator_list): | ||
| 568 | + return True | ||
| 569 | + for base in node.bases: | ||
| 570 | + name = _base_simple_name(base) | ||
| 571 | + parent = class_map.get(name) if name else None | ||
| 572 | + if parent is not None and id(parent) not in seen: | ||
| 573 | + stack.append(parent) | ||
| 574 | + return False | ||
| 575 | + | ||
| 576 | + | ||
| 577 | +def _last_field_line(cls_node: ast.ClassDef) -> int | None: | ||
| 578 | + """Last base line of the class's annotated fields (dataclass __init__ | ||
| 579 | + parameter order follows field order); None when the class has none.""" | ||
| 580 | + last = None | ||
| 581 | + for stmt in cls_node.body: | ||
| 582 | + if isinstance(stmt, ast.AnnAssign): | ||
| 583 | + end = stmt.end_lineno or stmt.lineno | ||
| 584 | + last = end if last is None else max(last, end) | ||
| 585 | + return last | ||
| 586 | + | ||
| 587 | + | ||
| 588 | +def _new_binding_insert_reason( | ||
| 589 | + add_texts, | ||
| 590 | + a: int, | ||
| 591 | + b: int, | ||
| 592 | + add_indent: int | None, | ||
| 593 | + ann_ctx: tuple | None, | ||
| 594 | +) -> str | None: | ||
| 595 | + """Skip reason when a pure insertion consists solely of assignments that | ||
| 596 | + bind provably-new names with inert right-hand sides; None to record. | ||
| 597 | + | ||
| 598 | + Base code cannot read a name that does not exist yet, so every reader of | ||
| 599 | + a first-time binding is itself an added line carrying its own anchor - | ||
| 600 | + the anchor here is noise (an import-time fingerprint in module/class | ||
| 601 | + scopes). Gates: no meaningful line ends in a comma (such lines are call | ||
| 602 | + or signature arguments, not statements); the block is all simple-name | ||
| 603 | + assignments; no target is a | ||
| 604 | + dunder or already occurs anywhere in the base file (Name occurrences, | ||
| 605 | + def/class/import/parameter bindings); right-hand sides are inert; and in | ||
| 606 | + class scope the class must be plain or dataclass-like, where plain | ||
| 607 | + assigns are never fields while annotated fields are only exempt when | ||
| 608 | + inserted at or after the last existing field (a middle field insertion | ||
| 609 | + would silently rebind positional constructor calls). | ||
| 610 | + """ | ||
| 611 | + if ann_ctx is None or not add_texts: | ||
| 612 | + return None | ||
| 613 | + effective = [t for t in add_texts if t.strip() and not t.strip().startswith("#")] | ||
| 614 | + if not effective or effective[-1].rstrip().endswith(","): | ||
| 615 | + return None | ||
| 616 | + scopes, class_map, has_future, known_names = ann_ctx | ||
| 617 | + stmts = _parse_assignments(add_texts) | ||
| 618 | + if not stmts: | ||
| 619 | + return None | ||
| 620 | + for stmt, names in stmts: | ||
| 621 | + if any(name.startswith("__") or name in known_names for name in names): | ||
| 622 | + return None | ||
| 623 | + if not _assign_rhs_lazy(stmt.value): | ||
| 624 | + return None | ||
| 625 | + scope = _insertion_scope(scopes, a, b, add_indent) | ||
| 626 | + if isinstance(scope, ast.ClassDef) and _class_consumes_annotations(scope, class_map): | ||
| 627 | + if not _class_dataclass_like(scope, class_map): | ||
| 628 | + return None | ||
| 629 | + if any(isinstance(stmt, ast.AnnAssign) for stmt, _ in stmts): | ||
| 630 | + last_field = _last_field_line(scope) | ||
| 631 | + if last_field is not None and a < last_field: | ||
| 632 | + return None | ||
| 633 | + return "first-time binding of new name(s)" | ||
| 634 | + | ||
| 635 | + | ||
| 449 | def _classify_candidate_pairs( | 636 | def _classify_candidate_pairs( |
| 450 | affected: set[int], | 637 | affected: set[int], |
| 451 | pairs: list[tuple], | 638 | pairs: list[tuple], |
| @@ -460,10 +647,17 @@ def _classify_candidate_pairs( | |||
| 460 | comment/docstring line in the base file AND the added lines are comments or | 647 | comment/docstring line in the base file AND the added lines are comments or |
| 461 | doc prose (not parseable Python), i.e. a pure comment/docstring change. | 648 | doc prose (not parseable Python), i.e. a pure comment/docstring change. |
| 462 | Pure type annotation changes are dropped in every direction (insertion, | 649 | Pure type annotation changes are dropped in every direction (insertion, |
| 463 | - replacement, deletion) when provably inert: function-local annotations are | 650 | + replacement, deletion) when provably inert: function-local annotations |
| 464 | - never evaluated; class-level ones need a plain class (no decorators, no | 651 | + are never evaluated; class-level ones need a plain class (no decorators, |
| 465 | - TypedDict/Protocol/Enum/BaseModel/NamedTuple base) and side-effect-free | 652 | + no TypedDict/Protocol/Enum/BaseModel/NamedTuple base) and side-effect-free |
| 466 | annotation expressions unless the file has future annotations. | 653 | annotation expressions unless the file has future annotations. |
| 654 | + First-time bindings are dropped when a pure insertion (no meaningful | ||
| 655 | + line ends in a comma - those are call/signature arguments) consists | ||
| 656 | + solely of assignments binding simple names that occur nowhere in the | ||
| 657 | + base file with inert right-hand sides (base code cannot read a name that | ||
| 658 | + does not exist yet; readers added in the same PR carry their own anchors); | ||
| 659 | + in dataclass-like classes annotated fields are only dropped when inserted | ||
| 660 | + at or after the last existing field. | ||
| 467 | Candidate pairs: without base content (or non-parseable Python) both sides | 661 | Candidate pairs: without base content (or non-parseable Python) both sides |
| 468 | of each pair are counted, bounded by the hunk. | 662 | of each pair are counted, bounded by the hunk. |
| 469 | """ | 663 | """ |
| @@ -532,6 +726,8 @@ def _classify_candidate_pairs( | |||
| 532 | reason = "belongs to a newly added function/class" | 726 | reason = "belongs to a newly added function/class" |
| 533 | if reason is None: | 727 | if reason is None: |
| 534 | reason = _pure_annotation_insert_reason(add_texts, a, b, add_indent, ann_ctx) | 728 | reason = _pure_annotation_insert_reason(add_texts, a, b, add_indent, ann_ctx) |
| 729 | + if reason is None: | ||
| 730 | + reason = _new_binding_insert_reason(add_texts, a, b, add_indent, ann_ctx) | ||
| 535 | # kind == 'blank': isolated blank deletion == one-line insertion | 731 | # kind == 'blank': isolated blank deletion == one-line insertion |
| 536 | elif _between_definitions(a, b, end_lines, start_lines, blanks): | 732 | elif _between_definitions(a, b, end_lines, start_lines, blanks): |
| 537 | reason = "between function/class definitions" | 733 | reason = "between function/class definitions" |
| @@ -11,7 +11,7 @@ __all__ = ["RepoAdapter", "VllmAscendAdapter", "SglangAdapter", "TorchNpuAdapter | |||
| 11 | _ADAPTERS = { | 11 | _ADAPTERS = { |
| 12 | "vllm_ascend": VllmAscendAdapter, | 12 | "vllm_ascend": VllmAscendAdapter, |
| 13 | "sglang": SglangAdapter, | 13 | "sglang": SglangAdapter, |
| 14 | - "torch_npu": TorchNpuAdapter, | 14 | + "pytorch": TorchNpuAdapter, |
| 15 | } | 15 | } |
| 16 | 16 | ||
| 17 | AVAILABLE_REPOS = tuple(_ADAPTERS.keys()) | 17 | AVAILABLE_REPOS = tuple(_ADAPTERS.keys()) |
| @@ -21,7 +21,7 @@ def get_adapter(repo_name: str) -> RepoAdapter: | |||
| 21 | """按仓库名获取适配器实例。 | 21 | """按仓库名获取适配器实例。 |
| 22 | 22 | ||
| 23 | Args: | 23 | Args: |
| 24 | - repo_name: 仓库名(vllm_ascend / sglang / torch_npu) | 24 | + repo_name: 仓库名(vllm_ascend / sglang / pytorch) |
| 25 | 25 | ||
| 26 | Returns: | 26 | Returns: |
| 27 | RepoAdapter 实例 | 27 | RepoAdapter 实例 |
| @@ -21,7 +21,7 @@ class RepoAdapter(ABC): | |||
| 21 | #: 覆盖率数据目录下的测试用例文件夹命名前缀;None 表示无固定前缀(vllm 用 tests__ 前缀 + cpu-ut) | 21 | #: 覆盖率数据目录下的测试用例文件夹命名前缀;None 表示无固定前缀(vllm 用 tests__ 前缀 + cpu-ut) |
| 22 | test_case_dir_prefix: str | None = None | 22 | test_case_dir_prefix: str | None = None |
| 23 | #: 覆盖率文件名 glob 模式(测试目录或其 covdata/ 子目录下)。 | 23 | #: 覆盖率文件名 glob 模式(测试目录或其 covdata/ 子目录下)。 |
| 24 | - #: 默认 "coverage*" 兼容 pytest-cov 裸文件(torch_npu: coverage)与原始格式 | 24 | + #: 默认 "coverage*" 兼容 pytest-cov 裸文件(pytorch: coverage)与原始格式 |
| 25 | #: 带后缀文件(vllm/sglang: coverage.linux-...-workflow.xxx)。 | 25 | #: 带后缀文件(vllm/sglang: coverage.linux-...-workflow.xxx)。 |
| 26 | coverage_file_glob: str = "coverage*" | 26 | coverage_file_glob: str = "coverage*" |
| 27 | 27 | ||
| @@ -1,14 +1,16 @@ | |||
| 1 | -"""torch_npu 仓库适配器(PyTorch / GitCode 社区)。 | 1 | +"""torch_npu 仓库适配器(检测上游 PyTorch:https://github.com/pytorch/pytorch)。 |
| 2 | 2 | ||
| 3 | -- REPO_NAME = "torch_npu",产品代码前缀 "torch_npu/" | 3 | +- REPO_NAME = "torch",产品代码前缀 "torch/",PR 源为 GitHub(--github-pr pytorch/pytorch#N) |
| 4 | -- 测试用例目录:目录名含 "__" 或以 "test_" 开头(如 test__xxx 编码路径), | 4 | +- 测试用例目录:目录名含 "__" 或以 "test_" 开头(如 _inductor__test_add), |
| 5 | 覆盖率文件位于 covdata/ 子目录 | 5 | 覆盖率文件位于 covdata/ 子目录 |
| 6 | - 测试文件规则:test/ 下 test_*.py | 6 | - 测试文件规则:test/ 下 test_*.py |
| 7 | - 测试名规范化:-- -> ::,__ -> /,文件级补 .py | 7 | - 测试名规范化:-- -> ::,__ -> /,文件级补 .py |
| 8 | - 粒度默认值跟随合并版:line=True, function=True | 8 | - 粒度默认值跟随合并版:line=True, function=True |
| 9 | -- 源文件解析使用基类默认候选(source_dir/torch_npu/<rel>,运行需 -s 指向 covstub) | 9 | +- 源文件解析使用基类默认候选(source_dir/torch/<rel>,运行需 -s 指向 covstub, |
| 10 | + covstub/torch/ 对应 pytorch/pytorch 仓库根目录) | ||
| 10 | """ | 11 | """ |
| 11 | 12 | ||
| 13 | +import re | ||
| 12 | from fnmatch import fnmatchcase | 14 | from fnmatch import fnmatchcase |
| 13 | from pathlib import PurePosixPath | 15 | from pathlib import PurePosixPath |
| 14 | 16 | ||
| @@ -18,10 +20,14 @@ from .base import RepoAdapter | |||
| 18 | class TorchNpuAdapter(RepoAdapter): | 20 | class TorchNpuAdapter(RepoAdapter): |
| 19 | """torch_npu 仓库适配器。""" | 21 | """torch_npu 仓库适配器。""" |
| 20 | 22 | ||
| 21 | - repo_name = "torch_npu" | 23 | + repo_name = "torch" |
| 22 | - product_code_prefix = "torch_npu/" | 24 | + product_code_prefix = "torch/" |
| 23 | test_case_dir_prefix = None | 25 | test_case_dir_prefix = None |
| 24 | 26 | ||
| 27 | + #: 精确匹配路径中独立的 torch/ 段(前为路径起点或 /), | ||
| 28 | + #: 避免 pytorch/、torchvision/、test_torch.py 等含 torch 子串的路径被误剥 | ||
| 29 | + _TORCH_PATH_RE = re.compile(r"(?:^|/)torch/(.+)") | ||
| 30 | + | ||
| 25 | #: 可纳入 PR 检测的测试根目录 | 31 | #: 可纳入 PR 检测的测试根目录 |
| 26 | TEST_ROOTS = ("test/",) | 32 | TEST_ROOTS = ("test/",) |
| 27 | 33 | ||
| @@ -30,7 +36,7 @@ class TorchNpuAdapter(RepoAdapter): | |||
| 30 | # ------------------------------------------------------------------ | 36 | # ------------------------------------------------------------------ |
| 31 | def is_test_case_dir(self, name: str) -> bool: | 37 | def is_test_case_dir(self, name: str) -> bool: |
| 32 | # PyTorch: "__" in name or name.startswith("test_") | 38 | # PyTorch: "__" in name or name.startswith("test_") |
| 33 | - return ("__" in name) or name.startswith("test_") | 39 | + return name.endswith(".py") or ("__" in name) or name.startswith("test_") |
| 34 | 40 | ||
| 35 | def normalize_test_name(self, test_name: str) -> str: | 41 | def normalize_test_name(self, test_name: str) -> str: |
| 36 | """ | 42 | """ |
| @@ -40,19 +46,16 @@ class TorchNpuAdapter(RepoAdapter): | |||
| 40 | """ | 46 | """ |
| 41 | # 先转换 -- 为 ::(函数级测试标记) | 47 | # 先转换 -- 为 ::(函数级测试标记) |
| 42 | result = test_name.replace("--", "::") | 48 | result = test_name.replace("--", "::") |
| 43 | - # 再转换 __ 为 / | 49 | + if result.endswith(".py") or (".py::" in result): |
| 44 | - result = result.replace("__", "/") | 50 | + if not result.startswith("test/"): |
| 45 | - # 无 ::(文件级测试)则补 .py 后缀 | 51 | + result = f"test/{result}" |
| 46 | - if "::" not in result: | 52 | + return result |
| 47 | - result = result + ".py" | ||
| 48 | return result | 53 | return result |
| 49 | 54 | ||
| 50 | def covered_path_to_rel(self, path: str) -> str | None: | 55 | def covered_path_to_rel(self, path: str) -> str | None: |
| 51 | - """覆盖率库路径包含 torch_npu/ 时剥离前缀得到相对路径,否则返回 None。""" | 56 | + """覆盖率库路径含独立的 torch/ 段时剥离前缀得到相对路径,否则返回 None。""" |
| 52 | - if self.repo_name not in path: | 57 | + match = self._TORCH_PATH_RE.search(path) |
| 53 | - return None | 58 | + return match.group(1) if match else None |
| 54 | - marker = f"{self.repo_name}/" | ||
| 55 | - return path.split(marker)[-1] if marker in path else path | ||
| 56 | 59 | ||
| 57 | # ------------------------------------------------------------------ | 60 | # ------------------------------------------------------------------ |
| 58 | # 测试文件检测(PR 内容检测) | 61 | # 测试文件检测(PR 内容检测) |
| @@ -3,7 +3,7 @@ const path = require("path"); | |||
| 3 | /** | 3 | /** |
| 4 | * Available repository adapters supported by the precision-test tool. | 4 | * Available repository adapters supported by the precision-test tool. |
| 5 | */ | 5 | */ |
| 6 | -const AVAILABLE_REPOS = ["vllm_ascend", "sglang", "torch_npu"]; | 6 | +const AVAILABLE_REPOS = ["vllm_ascend", "sglang", "pytorch"]; |
| 7 | 7 | ||
| 8 | /** | 8 | /** |
| 9 | * Whitelist of environment variable names safe to pass to the Python subprocess. | 9 | * Whitelist of environment variable names safe to pass to the Python subprocess. |