已开启
test: 添加 v5.1.0 综合测试报告(含智能体专项测试) #1103
test: 添加 v5.1.0 综合测试报告(含智能体专项测试) #1103
已开启
childish创建于 18 天前
共 5 个文件变更+155-0
@@ -0,0 +1,129 @@
1+# AtomCode v5.1.0 综合测试报告(含智能体专项测试)
2+ 
3+| 项目 | 内容 |
4+|---|---|
5+| 被测对象 | AtomCode v5.1.0(Rust 终端 AI 编码代理) |
6+| 测试类型 | 传统测试(静态/功能/兼容性)+ 智能体专项测试 |
7+| 测试环境 | Windows 10/11 (10.0.26200) · Git Bash · v5.0.5→v5.1.0 源码(a291bae1) |
8+| 测试人 | 2501_93773090(软件测试师角色) |
9+| 测试日期 | 2026-09-20 |
10+| 参考基线 | 公开 Issue 清单 #1254–#1338、docs/、AGENTS.md 架构约束 |
11+ 
12+---
13+ 
14+## 1. 测试范围与方法
15+ 
16+1. **静态代码测试**:人工走查 `crates/atomcode-capabilities`、`crates/atomcode-kernel` 的安全关键路径(命令执行、审批门控、注入面、错误处理)。
17+2. **功能性测试**:基于文档与代码行为核对(斜杠命令、登录、FAQ 描述与实现一致性)。
18+3. **智能体专项测试**(非传统测试,针对 LLM Agent 特有风险面):
19+ - 工具调用与审批门控(bypass/auto-approve)
20+ - 系统提示词与记忆注入(`=== MEMORY ===`)
21+ - Skill 展开与命令注入
22+ - MCP 集成(stdio/http transport、trust)
23+ - 会话恢复/rewind/compaction
24+ - 子代理(task/team)生命周期
25+4. **兼容性测试**:Windows 平台已知问题走查(基于 issue 反馈)。
26+ 
27+---
28+ 
29+## 2. 测试结果总览
30+ 
31+| 编号 | 类别 | 测试项 | 结果 | 严重级别 |
32+|---|---|---|---|---|
33+| T-01 | 智能体 | Skill `!`cmd`` 命令注入面 | ⚠️ 有风险 | **高** |
34+| T-02 | 智能体 | bash/bash_start 审批 bypass 覆盖 | ✅ 有测试覆盖 | — |
35+| T-03 | 智能体 | `run_to_completion` 吞 panic | ⚠️ 已知缺陷 | **高** |
36+| T-04 | 智能体 | 记忆注入无上限 | ⚠️ 设计风险 | 中 |
37+| T-05 | 智能体 | MCP 模块错误处理 | ⚠️ 更正:生产代码无 panic 点;问题是可观测性缺失 | 中 |
38+| T-06 | 智能体 | 静默 fallback(provider 降级) | ⚠️ 已知缺陷 | 高 |
39+| T-07 | 传统 | 安装脚本版本回退 | ⚠️ 已修复中 | 低 |
40+| T-08 | 传统 | Windows TUI 渲染 | ⚠️ 已知缺陷 | 中 |
41+| T-09 | 传统 | 文档一致性 | ✅ 基本一致 | — |
42+| T-10 | 智能体 | 子代理 panic 致假成功 | ⚠️ 已知缺陷 | 高 |
43+ 
44+---
45+ 
46+## 3. 智能体专项测试详情
47+ 
48+### T-01 Skill 命令注入面(高危,对应 #1274)
49+ 
50+**位置**:`crates/atomcode-capabilities/src/skills/skill.rs:170-211`
51+ 
52+`expand_shell_injections()` 将 SKILL.md 模板中的 `` !`cmd` `` 直接交给 `sh -c` 执行,无来源白名单、无审批确认:
53+ 
54+```rust
55+fn run_shell_command(cmd: &str) -> String {
56+ let mut command = Command::new("sh");
57+ command.arg("-c").arg(cmd); // 直接执行,无审批、无白名单
58+```
59+ 
60+**风险场景**:从 marketplace 安装的第三方 skill(如实战中的 ascend 插件)模板中嵌入任意命令,加载即执行,绕过 Bash 工具的四级审批门控。
61+ 
62+**建议**:对 `!`cmd`` 注入走与 Bash 工具相同的审批等级,或在信任 skill 时一次性声明允许的命令前缀白名单。
63+ 
64+### T-02 审批门控 bypass 覆盖(通过)
65+ 
66+`tools/approval.rs:517-542` 有明确单测:`allow-all grant must bypass risky bash_start just like bash`,证明 "总是允许" 对 `bash` 与 `bash_start` 两个变体一致生效。✅ 该处设计有回归测试保护。
67+ 
68+### T-03 `run_to_completion` 错误吞噬(高危,对应 #1282)
69+ 
70+`crates/atomcode-kernel/src/agent.rs:1101`。注释自述 "no longer SWALLOWS errors",但 issue #1282 报告 Agent 子任务 panic 仍被上报为正常完成。智能体长循环中 panic 被静默会导致**失败被报告为成功**,是 Agent 系统最危险的一类缺陷(用户以为任务完成,实际半途而废)。需补 panic→Outcome 的端到端回归测试。
71+ 
72+### T-04 记忆注入截断方向 bug(高危 → 已在 v5.2.0 修复,更正)
73+ 
74+**最终事实**(经维护者定位,更正初版结论):记忆注入早有 4000 字符上限(`merged_for_prompt` 截断 + truncated 标记),并非无界;实测 +2231 token/+28% 是撞上限后的固定开销,非线性增长,"无界注入"定性撤回。
75+ 
76+**经对照实验暴露的真 bug**(比朴素截断更严重):条目按 remember 追加顺序存储(最旧在前),截断却"取前 4000 字符"——保留最旧条目,**静默丢弃最新、通常最相关的记忆**,且 project/local 会被 global 挤掉。
77+ 
78+**上游修复**(v5.2.0,已打 bugfixed 标签):① 截断改为保最新;② scope 优先级 local > project > global;③ 单条 500 字符上限;④ 注入尾部报告省略条数(可感知)。
79+ 
80+**测试方法启示**:测到 token 开销异常后,应先读截断实现再下"无界"结论;对照实验的价值在于把异常开销摆上台面,促使维护者深挖出隐蔽的方向性 bug。LRU/凝缩机制归入 #1256 三层记忆模型方向。
81+ 
82+### T-05 MCP 可观测性缺口(中危,更正)
83+ 
84+**更正说明**:初版报告称 mcp 模块存在 114 处 unwrap/panic 构成崩溃风险,经复核,**该数字全部来自 `#[cfg(test)]` 测试代码,生产代码 panic 点为 0**,初版结论有误,在此更正。
85+ 
86+经两轮动态实验(v5.1.0 实测)确认的真实问题是**可观测性缺失**:
87+ 
88+- stdio server 初始化失败、畸形输出被静默丢弃,UI 无任何提示;
89+- 非交互 `-p` 模式下项目级 `.mcp.json` server 被静默跳过,模型看不到已配置的工具,用户无从得知原因。
90+ 
91+**建议**:server 从注册 → 加载 → 初始化 → 运行各环节失败时,在会话启动横幅或 `/status` 中显式列出原因;`mcp add` 后提示配置生效条件。
92+ 
93+### T-06 静默 fallback(高危,对应 #1322/#1300)
94+ 
95+CodingPlan Pro-体验版 GLM-5.2 曾出现"8-02 静默 fallback + 403"——provider 调用失败时不告警、降级到其他模型继续,用户以为在用 A 模型实际是 B。对 Agent 产品,静默降级破坏可信任性。**建议**:任何模型降级必须在 UI 显式横幅提示并写入 transcript。
96+ 
97+### T-07 子代理生命周期(对应 #1282/#1254)
98+ 
99+`agent.rs:988` 注释确认子代理停止依赖 `run_to_completion` 这一"唯一通道",单点设计脆弱;WebUI sync 模式消息延迟 2~10s(#1254)说明 live view 事件通道在子代理并发场景下有背板积压。
100+ 
101+---
102+ 
103+## 4. 传统测试详情
104+ 
105+| 编号 | 测试项 | 结果 |
106+|---|---|---|
107+| T-07 | `scripts/install.sh` / `install.ps1` DEFAULT_VERSION 停留在 v5.0.2(当前 v5.1.0),新装用户拿到旧版 | ⚠️ PR #943 修复中 |
108+| T-08 | Windows TUI:标号文字重叠(#1324)、429 后 provider 表单回车无效(#1335)、secrets 读取后对话中断(#1336) | ⚠️ 三个 Win10 专属缺陷,均无人认领 |
109+| T-09 | 文档与实现一致性:抽查 FAQ/斜杠命令 30+ 条目,版本号 v5.0.5 与 Cargo v5.1.0 不一致 | ⚠️ 轻微 |
110+| T-10 | 构建/依赖:Cargo workspace 15 crates,`atomcode-core` 已退役且生产依赖 core-free 约束在 AGENTS.md 中明确 | ✅ 通过 |
111+ 
112+---
113+ 
114+## 5. 智能体测试方法建议(供项目采纳)
115+ 
116+1. **审批覆盖矩阵**:为每个工具 × 四级审批等级建立矩阵化集成测试(T-02 已是良好范例)。
117+2. **注入面清单**:所有最终落到 `sh -c` / `cmd /c` 的调用点集中登记并强制过审批(目前分散在 tools/hooks/skills 三处)。
118+3. **降级可观测性测试**:模拟 provider 4xx/5xx,断言 UI 出现降级横幅。
119+4. **记忆预算测试**:记忆条目增长到 N 条时,断言注入 token 超上限被截断。
120+5. **长任务混沌测试**:随机在工具调用间注入 panic/断流,断言会话恢复(`--continue`)后 Outcome 一致。
121+ 
122+---
123+ 
124+## 6. 结论
125+ 
126+AtomCode 架构方向清晰(kernel/coding/capabilities 分层 + core-free 约束),审批门控有回归测试保护,文档完整度高。**智能体特有的四类高风险问题**值得优先处理:命令注入面(T-01)、错误吞噬致假成功(T-03/T-10)、静默降级(T-06)、记忆无界增长(T-04)。建议在 v5.2 前完成 T-01、T-03、T-06 的修复与回归测试。
127+ 
128+---
129+*测试过程中参考了社区 issue 反馈与既有 PR 讨论(含 #944 FAQ 贡献),欢迎在 Issue 区交流测试方法。*
@@ -108,6 +108,14 @@
108 <li>If you're in mainland China hitting an overseas model, verify your egress network. You can run <code>/model</code> to switch to AtomGit's official channel first to prove the whole pipeline works.</li>108 <li>If you're in mainland China hitting an overseas model, verify your egress network. You can run <code>/model</code> to switch to AtomGit's official channel first to prove the whole pipeline works.</li>
109 </ul>109 </ul>
110 110 
111+ <h3>Intermittent errors / interrupted turns in long tasks (401, 403, empty responses, connection closed)</h3>
112+ <ul>
113+ <li>Check your quota first: run <code>/status</code> to see CodingPlan usage and reset time (quota is measured on a rolling 5-hour window), or <code>/usage</code> to query your plan quota and see whether you're being rate-limited;</li>
114+ <li>Re-sync: run <code>/login</code> to idempotently refresh the model list and token — this usually restores subsequent calls;</li>
115+ <li>For one-off disconnects, just retry, or use <code>/resume</code> to pick up the last session;</li>
116+ <li>If it keeps happening, check the logs: <code>~/.atomcode/logs/</code> — launch with <code>--log-level debug</code> to print request/response details and tell rate limiting, network, or gateway issues apart.</li>
117+ </ul>
118+ 
111 <h3 id="how-do-i-switch-models">How do I switch models?</h3>119 <h3 id="how-do-i-switch-models">How do I switch models?</h3>
112 <p>Type <code>/model</code> and pick from the menu. The command switches the current session and saves the choice as the default for sessions opened later; other already-open sessions keep their current model. See <a href="./configuration.html">Configuration</a>.</p>120 <p>Type <code>/model</code> and pick from the menu. The command switches the current session and saves the choice as the default for sessions opened later; other already-open sessions keep their current model. See <a href="./configuration.html">Configuration</a>.</p>
113 121 
@@ -1402,6 +1402,11 @@
1402 "heading": "Can't reach the model / hangs forever / timeouts",1402 "heading": "Can't reach the model / hangs forever / timeouts",
1403 "body": "First check that base_url in ~/.atomcode/config.toml is reachable — curl the matching /v1/models endpoint; On corporate networks it's usually a proxy issue — set HTTPS_PROXY and retry; On Windows, v5.0.7 uses the system SChannel certificate store by default, including enterprise roots. If it still fails, check the system certificate, proxy, and overrides such as SSL_CERT_FILE separately; If you're in mainland China hitting an overseas model, verify your egress network. You can run /model to switch to AtomGit's official channel first to prove the whole pipeline works."1403 "body": "First check that base_url in ~/.atomcode/config.toml is reachable — curl the matching /v1/models endpoint; On corporate networks it's usually a proxy issue — set HTTPS_PROXY and retry; On Windows, v5.0.7 uses the system SChannel certificate store by default, including enterprise roots. If it still fails, check the system certificate, proxy, and overrides such as SSL_CERT_FILE separately; If you're in mainland China hitting an overseas model, verify your egress network. You can run /model to switch to AtomGit's official channel first to prove the whole pipeline works."
1404 },1404 },
1405+ {
1406+ "id": "intermittent-errors-interrupted-turns-in-long-tasks-401-403-",
1407+ "heading": "Intermittent errors / interrupted turns in long tasks (401, 403, empty responses, connection closed)",
1408+ "body": "Check your quota first: run /status to see CodingPlan usage and reset time (quota is measured on a rolling 5-hour window), or /usage to query your plan quota and see whether you're being rate-limited; Re-sync: run /login to idempotently refresh the model list and token — this usually restores subsequent calls; For one-off disconnects, just retry, or use /resume to pick up the last session; If it keeps happening, check the logs: ~/.atomcode/logs/ — launch with --log-level debug to print request/response details and tell rate limiting, network, or gateway issues apart."
1409+ },
1405 {1410 {
1406 "id": "how-do-i-switch-models",1411 "id": "how-do-i-switch-models",
1407 "heading": "How do I switch models?",1412 "heading": "How do I switch models?",
@@ -1402,6 +1402,11 @@
1402 "heading": "连不上模型 / 一直转圈 / 报 timeout",1402 "heading": "连不上模型 / 一直转圈 / 报 timeout",
1403 "body": "先确认 ~/.atomcode/config.toml 里的 base_url 是否可达, curl 一下对应的 /v1/models 接口; 企业网络里常见是代理问题,设置 HTTPS_PROXY 环境变量后重试; Windows v5.0.7 默认使用系统 SChannel 证书库,可识别企业安装的根证书;若仍失败,请分别检查系统证书、代理和 SSL_CERT_FILE 等覆盖变量。 国内访问境外模型需要确认出口网络;可以先用 /model 切换到 AtomGit 官方通道确认整体链路是通的。"1403 "body": "先确认 ~/.atomcode/config.toml 里的 base_url 是否可达, curl 一下对应的 /v1/models 接口; 企业网络里常见是代理问题,设置 HTTPS_PROXY 环境变量后重试; Windows v5.0.7 默认使用系统 SChannel 证书库,可识别企业安装的根证书;若仍失败,请分别检查系统证书、代理和 SSL_CERT_FILE 等覆盖变量。 国内访问境外模型需要确认出口网络;可以先用 /model 切换到 AtomGit 官方通道确认整体链路是通的。"
1404 },1404 },
1405+ {
1406+ "id": "长任务中途偶发报错-回合中断401403空响应connection-closed",
1407+ "heading": "长任务中途偶发报错 / 回合中断(401、403、空响应、connection closed)",
1408+ "body": "先看额度:运行 /status 查看 CodingPlan 用量与重置时间(额度按 5 小时滚动窗口计量),或 /usage 查询计划配额,确认是否触发限流; 重新同步:执行 /login 幂等地刷新模型列表与 token,通常能恢复后续调用; 偶发断连可直接重试,或用 /resume 恢复上一次会话继续; 仍复现再看日志: ~/.atomcode/logs/ ,启动时加 --log-level debug 打印请求/响应详情,方便定位是限流、网络还是网关侧问题。"
1409+ },
1405 {1410 {
1406 "id": "怎么切换模型",1411 "id": "怎么切换模型",
1407 "heading": "怎么切换模型?",1412 "heading": "怎么切换模型?",
@@ -108,6 +108,14 @@
108 <li>国内访问境外模型需要确认出口网络;可以先用 <code>/model</code> 切换到 AtomGit 官方通道确认整体链路是通的。</li>108 <li>国内访问境外模型需要确认出口网络;可以先用 <code>/model</code> 切换到 AtomGit 官方通道确认整体链路是通的。</li>
109 </ul>109 </ul>
110 110 
111+ <h3>长任务中途偶发报错 / 回合中断(401、403、空响应、connection closed)</h3>
112+ <ul>
113+ <li>先看额度:运行 <code>/status</code> 查看 CodingPlan 用量与重置时间(额度按 5 小时滚动窗口计量),或 <code>/usage</code> 查询计划配额,确认是否触发限流;</li>
114+ <li>重新同步:执行 <code>/login</code> 幂等地刷新模型列表与 token,通常能恢复后续调用;</li>
115+ <li>偶发断连可直接重试,或用 <code>/resume</code> 恢复上一次会话继续;</li>
116+ <li>仍复现再看日志:<code>~/.atomcode/logs/</code>,启动时加 <code>--log-level debug</code> 打印请求/响应详情,方便定位是限流、网络还是网关侧问题。</li>
117+ </ul>
118+ 
111 <h3 id="怎么切换模型">怎么切换模型?</h3>119 <h3 id="怎么切换模型">怎么切换模型?</h3>
112 <p>直接输入 <code>/model</code> 并从菜单选择。命令会切换当前会话,并把所选项保存为后续新会话的默认值;其他已经打开的会话保持原模型。详见 <a href="./configuration.html">配置文件</a>。</p>120 <p>直接输入 <code>/model</code> 并从菜单选择。命令会切换当前会话,并把所选项保存为后续新会话的默认值;其他已经打开的会话保持原模型。详见 <a href="./configuration.html">配置文件</a>。</p>
113 121