已开启
test: 添加 v5.1.0 综合测试报告(含智能体专项测试) #1103
childish创建于 18 天前
test: 添加 v5.1.0 综合测试报告(含智能体专项测试) #1103
已开启
共 5 个文件变更+155-0
| @@ -0,0 +1,129 @@ | |||
| 1 | +# AtomCode v5.1.0 综合测试报告(含智能体专项测试) | ||
| 2 | + | ||
| 3 | +| 项目 | 内容 | | ||
| 4 | +|---|---| | ||
| 5 | +| 被测对象 | AtomCode v5.1.0(Rust 终端 AI 编码代理) | | ||
| 6 | +| 测试类型 | 传统测试(静态/功能/兼容性)+ 智能体专项测试 | | ||
| 7 | +| 测试环境 | Windows 10/11 (10.0.26200) · Git Bash · v5.0.5→v5.1.0 源码(a291bae1) | | ||
| 8 | +| 测试人 | 2501_93773090(软件测试师角色) | | ||
| 9 | +| 测试日期 | 2026-09-20 | | ||
| 10 | +| 参考基线 | 公开 Issue 清单 #1254–#1338、docs/、AGENTS.md 架构约束 | | ||
| 11 | + | ||
| 12 | +--- | ||
| 13 | + | ||
| 14 | +## 1. 测试范围与方法 | ||
| 15 | + | ||
| 16 | +1. **静态代码测试**:人工走查 `crates/atomcode-capabilities`、`crates/atomcode-kernel` 的安全关键路径(命令执行、审批门控、注入面、错误处理)。 | ||
| 17 | +2. **功能性测试**:基于文档与代码行为核对(斜杠命令、登录、FAQ 描述与实现一致性)。 | ||
| 18 | +3. **智能体专项测试**(非传统测试,针对 LLM Agent 特有风险面): | ||
| 19 | + - 工具调用与审批门控(bypass/auto-approve) | ||
| 20 | + - 系统提示词与记忆注入(`=== MEMORY ===`) | ||
| 21 | + - Skill 展开与命令注入 | ||
| 22 | + - MCP 集成(stdio/http transport、trust) | ||
| 23 | + - 会话恢复/rewind/compaction | ||
| 24 | + - 子代理(task/team)生命周期 | ||
| 25 | +4. **兼容性测试**:Windows 平台已知问题走查(基于 issue 反馈)。 | ||
| 26 | + | ||
| 27 | +--- | ||
| 28 | + | ||
| 29 | +## 2. 测试结果总览 | ||
| 30 | + | ||
| 31 | +| 编号 | 类别 | 测试项 | 结果 | 严重级别 | | ||
| 32 | +|---|---|---|---|---| | ||
| 33 | +| T-01 | 智能体 | Skill `!`cmd`` 命令注入面 | ⚠️ 有风险 | **高** | | ||
| 34 | +| T-02 | 智能体 | bash/bash_start 审批 bypass 覆盖 | ✅ 有测试覆盖 | — | | ||
| 35 | +| T-03 | 智能体 | `run_to_completion` 吞 panic | ⚠️ 已知缺陷 | **高** | | ||
| 36 | +| T-04 | 智能体 | 记忆注入无上限 | ⚠️ 设计风险 | 中 | | ||
| 37 | +| T-05 | 智能体 | MCP 模块错误处理 | ⚠️ 更正:生产代码无 panic 点;问题是可观测性缺失 | 中 | | ||
| 38 | +| T-06 | 智能体 | 静默 fallback(provider 降级) | ⚠️ 已知缺陷 | 高 | | ||
| 39 | +| T-07 | 传统 | 安装脚本版本回退 | ⚠️ 已修复中 | 低 | | ||
| 40 | +| T-08 | 传统 | Windows TUI 渲染 | ⚠️ 已知缺陷 | 中 | | ||
| 41 | +| T-09 | 传统 | 文档一致性 | ✅ 基本一致 | — | | ||
| 42 | +| T-10 | 智能体 | 子代理 panic 致假成功 | ⚠️ 已知缺陷 | 高 | | ||
| 43 | + | ||
| 44 | +--- | ||
| 45 | + | ||
| 46 | +## 3. 智能体专项测试详情 | ||
| 47 | + | ||
| 48 | +### T-01 Skill 命令注入面(高危,对应 #1274) | ||
| 49 | + | ||
| 50 | +**位置**:`crates/atomcode-capabilities/src/skills/skill.rs:170-211` | ||
| 51 | + | ||
| 52 | +`expand_shell_injections()` 将 SKILL.md 模板中的 `` !`cmd` `` 直接交给 `sh -c` 执行,无来源白名单、无审批确认: | ||
| 53 | + | ||
| 54 | +```rust | ||
| 55 | +fn run_shell_command(cmd: &str) -> String { | ||
| 56 | + let mut command = Command::new("sh"); | ||
| 57 | + command.arg("-c").arg(cmd); // 直接执行,无审批、无白名单 | ||
| 58 | +``` | ||
| 59 | + | ||
| 60 | +**风险场景**:从 marketplace 安装的第三方 skill(如实战中的 ascend 插件)模板中嵌入任意命令,加载即执行,绕过 Bash 工具的四级审批门控。 | ||
| 61 | + | ||
| 62 | +**建议**:对 `!`cmd`` 注入走与 Bash 工具相同的审批等级,或在信任 skill 时一次性声明允许的命令前缀白名单。 | ||
| 63 | + | ||
| 64 | +### T-02 审批门控 bypass 覆盖(通过) | ||
| 65 | + | ||
| 66 | +`tools/approval.rs:517-542` 有明确单测:`allow-all grant must bypass risky bash_start just like bash`,证明 "总是允许" 对 `bash` 与 `bash_start` 两个变体一致生效。✅ 该处设计有回归测试保护。 | ||
| 67 | + | ||
| 68 | +### T-03 `run_to_completion` 错误吞噬(高危,对应 #1282) | ||
| 69 | + | ||
| 70 | +`crates/atomcode-kernel/src/agent.rs:1101`。注释自述 "no longer SWALLOWS errors",但 issue #1282 报告 Agent 子任务 panic 仍被上报为正常完成。智能体长循环中 panic 被静默会导致**失败被报告为成功**,是 Agent 系统最危险的一类缺陷(用户以为任务完成,实际半途而废)。需补 panic→Outcome 的端到端回归测试。 | ||
| 71 | + | ||
| 72 | +### T-04 记忆注入截断方向 bug(高危 → 已在 v5.2.0 修复,更正) | ||
| 73 | + | ||
| 74 | +**最终事实**(经维护者定位,更正初版结论):记忆注入早有 4000 字符上限(`merged_for_prompt` 截断 + truncated 标记),并非无界;实测 +2231 token/+28% 是撞上限后的固定开销,非线性增长,"无界注入"定性撤回。 | ||
| 75 | + | ||
| 76 | +**经对照实验暴露的真 bug**(比朴素截断更严重):条目按 remember 追加顺序存储(最旧在前),截断却"取前 4000 字符"——保留最旧条目,**静默丢弃最新、通常最相关的记忆**,且 project/local 会被 global 挤掉。 | ||
| 77 | + | ||
| 78 | +**上游修复**(v5.2.0,已打 bugfixed 标签):① 截断改为保最新;② scope 优先级 local > project > global;③ 单条 500 字符上限;④ 注入尾部报告省略条数(可感知)。 | ||
| 79 | + | ||
| 80 | +**测试方法启示**:测到 token 开销异常后,应先读截断实现再下"无界"结论;对照实验的价值在于把异常开销摆上台面,促使维护者深挖出隐蔽的方向性 bug。LRU/凝缩机制归入 #1256 三层记忆模型方向。 | ||
| 81 | + | ||
| 82 | +### T-05 MCP 可观测性缺口(中危,更正) | ||
| 83 | + | ||
| 84 | +**更正说明**:初版报告称 mcp 模块存在 114 处 unwrap/panic 构成崩溃风险,经复核,**该数字全部来自 `#[cfg(test)]` 测试代码,生产代码 panic 点为 0**,初版结论有误,在此更正。 | ||
| 85 | + | ||
| 86 | +经两轮动态实验(v5.1.0 实测)确认的真实问题是**可观测性缺失**: | ||
| 87 | + | ||
| 88 | +- stdio server 初始化失败、畸形输出被静默丢弃,UI 无任何提示; | ||
| 89 | +- 非交互 `-p` 模式下项目级 `.mcp.json` server 被静默跳过,模型看不到已配置的工具,用户无从得知原因。 | ||
| 90 | + | ||
| 91 | +**建议**:server 从注册 → 加载 → 初始化 → 运行各环节失败时,在会话启动横幅或 `/status` 中显式列出原因;`mcp add` 后提示配置生效条件。 | ||
| 92 | + | ||
| 93 | +### T-06 静默 fallback(高危,对应 #1322/#1300) | ||
| 94 | + | ||
| 95 | +CodingPlan Pro-体验版 GLM-5.2 曾出现"8-02 静默 fallback + 403"——provider 调用失败时不告警、降级到其他模型继续,用户以为在用 A 模型实际是 B。对 Agent 产品,静默降级破坏可信任性。**建议**:任何模型降级必须在 UI 显式横幅提示并写入 transcript。 | ||
| 96 | + | ||
| 97 | +### T-07 子代理生命周期(对应 #1282/#1254) | ||
| 98 | + | ||
| 99 | +`agent.rs:988` 注释确认子代理停止依赖 `run_to_completion` 这一"唯一通道",单点设计脆弱;WebUI sync 模式消息延迟 2~10s(#1254)说明 live view 事件通道在子代理并发场景下有背板积压。 | ||
| 100 | + | ||
| 101 | +--- | ||
| 102 | + | ||
| 103 | +## 4. 传统测试详情 | ||
| 104 | + | ||
| 105 | +| 编号 | 测试项 | 结果 | | ||
| 106 | +|---|---|---| | ||
| 107 | +| T-07 | `scripts/install.sh` / `install.ps1` DEFAULT_VERSION 停留在 v5.0.2(当前 v5.1.0),新装用户拿到旧版 | ⚠️ PR #943 修复中 | | ||
| 108 | +| T-08 | Windows TUI:标号文字重叠(#1324)、429 后 provider 表单回车无效(#1335)、secrets 读取后对话中断(#1336) | ⚠️ 三个 Win10 专属缺陷,均无人认领 | | ||
| 109 | +| T-09 | 文档与实现一致性:抽查 FAQ/斜杠命令 30+ 条目,版本号 v5.0.5 与 Cargo v5.1.0 不一致 | ⚠️ 轻微 | | ||
| 110 | +| T-10 | 构建/依赖:Cargo workspace 15 crates,`atomcode-core` 已退役且生产依赖 core-free 约束在 AGENTS.md 中明确 | ✅ 通过 | | ||
| 111 | + | ||
| 112 | +--- | ||
| 113 | + | ||
| 114 | +## 5. 智能体测试方法建议(供项目采纳) | ||
| 115 | + | ||
| 116 | +1. **审批覆盖矩阵**:为每个工具 × 四级审批等级建立矩阵化集成测试(T-02 已是良好范例)。 | ||
| 117 | +2. **注入面清单**:所有最终落到 `sh -c` / `cmd /c` 的调用点集中登记并强制过审批(目前分散在 tools/hooks/skills 三处)。 | ||
| 118 | +3. **降级可观测性测试**:模拟 provider 4xx/5xx,断言 UI 出现降级横幅。 | ||
| 119 | +4. **记忆预算测试**:记忆条目增长到 N 条时,断言注入 token 超上限被截断。 | ||
| 120 | +5. **长任务混沌测试**:随机在工具调用间注入 panic/断流,断言会话恢复(`--continue`)后 Outcome 一致。 | ||
| 121 | + | ||
| 122 | +--- | ||
| 123 | + | ||
| 124 | +## 6. 结论 | ||
| 125 | + | ||
| 126 | +AtomCode 架构方向清晰(kernel/coding/capabilities 分层 + core-free 约束),审批门控有回归测试保护,文档完整度高。**智能体特有的四类高风险问题**值得优先处理:命令注入面(T-01)、错误吞噬致假成功(T-03/T-10)、静默降级(T-06)、记忆无界增长(T-04)。建议在 v5.2 前完成 T-01、T-03、T-06 的修复与回归测试。 | ||
| 127 | + | ||
| 128 | +--- | ||
| 129 | +*测试过程中参考了社区 issue 反馈与既有 PR 讨论(含 #944 FAQ 贡献),欢迎在 Issue 区交流测试方法。* | ||
| @@ -108,6 +108,14 @@ | |||
| 108 | <li>If you're in mainland China hitting an overseas model, verify your egress network. You can run <code>/model</code> to switch to AtomGit's official channel first to prove the whole pipeline works.</li> | 108 | <li>If you're in mainland China hitting an overseas model, verify your egress network. You can run <code>/model</code> to switch to AtomGit's official channel first to prove the whole pipeline works.</li> |
| 109 | </ul> | 109 | </ul> |
| 110 | 110 | ||
| 111 | + <h3>Intermittent errors / interrupted turns in long tasks (401, 403, empty responses, connection closed)</h3> | ||
| 112 | + <ul> | ||
| 113 | + <li>Check your quota first: run <code>/status</code> to see CodingPlan usage and reset time (quota is measured on a rolling 5-hour window), or <code>/usage</code> to query your plan quota and see whether you're being rate-limited;</li> | ||
| 114 | + <li>Re-sync: run <code>/login</code> to idempotently refresh the model list and token — this usually restores subsequent calls;</li> | ||
| 115 | + <li>For one-off disconnects, just retry, or use <code>/resume</code> to pick up the last session;</li> | ||
| 116 | + <li>If it keeps happening, check the logs: <code>~/.atomcode/logs/</code> — launch with <code>--log-level debug</code> to print request/response details and tell rate limiting, network, or gateway issues apart.</li> | ||
| 117 | + </ul> | ||
| 118 | + | ||
| 111 | <h3 id="how-do-i-switch-models">How do I switch models?</h3> | 119 | <h3 id="how-do-i-switch-models">How do I switch models?</h3> |
| 112 | <p>Type <code>/model</code> and pick from the menu. The command switches the current session and saves the choice as the default for sessions opened later; other already-open sessions keep their current model. See <a href="./configuration.html">Configuration</a>.</p> | 120 | <p>Type <code>/model</code> and pick from the menu. The command switches the current session and saves the choice as the default for sessions opened later; other already-open sessions keep their current model. See <a href="./configuration.html">Configuration</a>.</p> |
| 113 | 121 | ||
| @@ -1402,6 +1402,11 @@ | |||
| 1402 | "heading": "Can't reach the model / hangs forever / timeouts", | 1402 | "heading": "Can't reach the model / hangs forever / timeouts", |
| 1403 | "body": "First check that base_url in ~/.atomcode/config.toml is reachable — curl the matching /v1/models endpoint; On corporate networks it's usually a proxy issue — set HTTPS_PROXY and retry; On Windows, v5.0.7 uses the system SChannel certificate store by default, including enterprise roots. If it still fails, check the system certificate, proxy, and overrides such as SSL_CERT_FILE separately; If you're in mainland China hitting an overseas model, verify your egress network. You can run /model to switch to AtomGit's official channel first to prove the whole pipeline works." | 1403 | "body": "First check that base_url in ~/.atomcode/config.toml is reachable — curl the matching /v1/models endpoint; On corporate networks it's usually a proxy issue — set HTTPS_PROXY and retry; On Windows, v5.0.7 uses the system SChannel certificate store by default, including enterprise roots. If it still fails, check the system certificate, proxy, and overrides such as SSL_CERT_FILE separately; If you're in mainland China hitting an overseas model, verify your egress network. You can run /model to switch to AtomGit's official channel first to prove the whole pipeline works." |
| 1404 | }, | 1404 | }, |
| 1405 | + { | ||
| 1406 | + "id": "intermittent-errors-interrupted-turns-in-long-tasks-401-403-", | ||
| 1407 | + "heading": "Intermittent errors / interrupted turns in long tasks (401, 403, empty responses, connection closed)", | ||
| 1408 | + "body": "Check your quota first: run /status to see CodingPlan usage and reset time (quota is measured on a rolling 5-hour window), or /usage to query your plan quota and see whether you're being rate-limited; Re-sync: run /login to idempotently refresh the model list and token — this usually restores subsequent calls; For one-off disconnects, just retry, or use /resume to pick up the last session; If it keeps happening, check the logs: ~/.atomcode/logs/ — launch with --log-level debug to print request/response details and tell rate limiting, network, or gateway issues apart." | ||
| 1409 | + }, | ||
| 1405 | { | 1410 | { |
| 1406 | "id": "how-do-i-switch-models", | 1411 | "id": "how-do-i-switch-models", |
| 1407 | "heading": "How do I switch models?", | 1412 | "heading": "How do I switch models?", |
| @@ -1402,6 +1402,11 @@ | |||
| 1402 | "heading": "连不上模型 / 一直转圈 / 报 timeout", | 1402 | "heading": "连不上模型 / 一直转圈 / 报 timeout", |
| 1403 | "body": "先确认 ~/.atomcode/config.toml 里的 base_url 是否可达, curl 一下对应的 /v1/models 接口; 企业网络里常见是代理问题,设置 HTTPS_PROXY 环境变量后重试; Windows v5.0.7 默认使用系统 SChannel 证书库,可识别企业安装的根证书;若仍失败,请分别检查系统证书、代理和 SSL_CERT_FILE 等覆盖变量。 国内访问境外模型需要确认出口网络;可以先用 /model 切换到 AtomGit 官方通道确认整体链路是通的。" | 1403 | "body": "先确认 ~/.atomcode/config.toml 里的 base_url 是否可达, curl 一下对应的 /v1/models 接口; 企业网络里常见是代理问题,设置 HTTPS_PROXY 环境变量后重试; Windows v5.0.7 默认使用系统 SChannel 证书库,可识别企业安装的根证书;若仍失败,请分别检查系统证书、代理和 SSL_CERT_FILE 等覆盖变量。 国内访问境外模型需要确认出口网络;可以先用 /model 切换到 AtomGit 官方通道确认整体链路是通的。" |
| 1404 | }, | 1404 | }, |
| 1405 | + { | ||
| 1406 | + "id": "长任务中途偶发报错-回合中断401403空响应connection-closed", | ||
| 1407 | + "heading": "长任务中途偶发报错 / 回合中断(401、403、空响应、connection closed)", | ||
| 1408 | + "body": "先看额度:运行 /status 查看 CodingPlan 用量与重置时间(额度按 5 小时滚动窗口计量),或 /usage 查询计划配额,确认是否触发限流; 重新同步:执行 /login 幂等地刷新模型列表与 token,通常能恢复后续调用; 偶发断连可直接重试,或用 /resume 恢复上一次会话继续; 仍复现再看日志: ~/.atomcode/logs/ ,启动时加 --log-level debug 打印请求/响应详情,方便定位是限流、网络还是网关侧问题。" | ||
| 1409 | + }, | ||
| 1405 | { | 1410 | { |
| 1406 | "id": "怎么切换模型", | 1411 | "id": "怎么切换模型", |
| 1407 | "heading": "怎么切换模型?", | 1412 | "heading": "怎么切换模型?", |
| @@ -108,6 +108,14 @@ | |||
| 108 | <li>国内访问境外模型需要确认出口网络;可以先用 <code>/model</code> 切换到 AtomGit 官方通道确认整体链路是通的。</li> | 108 | <li>国内访问境外模型需要确认出口网络;可以先用 <code>/model</code> 切换到 AtomGit 官方通道确认整体链路是通的。</li> |
| 109 | </ul> | 109 | </ul> |
| 110 | 110 | ||
| 111 | + <h3>长任务中途偶发报错 / 回合中断(401、403、空响应、connection closed)</h3> | ||
| 112 | + <ul> | ||
| 113 | + <li>先看额度:运行 <code>/status</code> 查看 CodingPlan 用量与重置时间(额度按 5 小时滚动窗口计量),或 <code>/usage</code> 查询计划配额,确认是否触发限流;</li> | ||
| 114 | + <li>重新同步:执行 <code>/login</code> 幂等地刷新模型列表与 token,通常能恢复后续调用;</li> | ||
| 115 | + <li>偶发断连可直接重试,或用 <code>/resume</code> 恢复上一次会话继续;</li> | ||
| 116 | + <li>仍复现再看日志:<code>~/.atomcode/logs/</code>,启动时加 <code>--log-level debug</code> 打印请求/响应详情,方便定位是限流、网络还是网关侧问题。</li> | ||
| 117 | + </ul> | ||
| 118 | + | ||
| 111 | <h3 id="怎么切换模型">怎么切换模型?</h3> | 119 | <h3 id="怎么切换模型">怎么切换模型?</h3> |
| 112 | <p>直接输入 <code>/model</code> 并从菜单选择。命令会切换当前会话,并把所选项保存为后续新会话的默认值;其他已经打开的会话保持原模型。详见 <a href="./configuration.html">配置文件</a>。</p> | 120 | <p>直接输入 <code>/model</code> 并从菜单选择。命令会切换当前会话,并把所选项保存为后续新会话的默认值;其他已经打开的会话保持原模型。详见 <a href="./configuration.html">配置文件</a>。</p> |
| 113 | 121 | ||