| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
fix(evals): stop scoring crashing on, and inventing counts from, recorded data (#225) tools/evals/score.py documents itself as scoring "without loading files or deriving missing observations", and aggregate() promises to "never estimate missing usage". Two things broke that contract. 1. opens.index(target) was called unguarded. It is only reached when route_correct and answer_correct are both true -- but route_correct is only DERIVED from opens when the harness did not record it. A harness that records route_correct itself, while opens does not contain the target verbatim, hit ValueError: opens=["chapters/ch01.md"] target="chapters/ch02.md" -> ValueError opens=[] target="a.md" -> ValueError opens=["./chapters/ch02.md"] target="chapters/ch02.md" -> ValueError score() maps over every trajectory, so one such row aborted the whole scoring run rather than one question. The position is now computed once, guarded by membership, and absence simply means there is no evidence of irrelevant opens before the target. 2. isinstance(value, int) accepted True, because bool subclasses int in Python. A JSON `true` in a usage field was treated as a recorded count and summed as 1 by aggregate() -- exactly the estimate the module promises not to make. _count() now rejects bool explicitly. Derived routing is unchanged: when the harness records nothing, routing is still derived from opens, and target-after-other-opens is still classified irrelevant_opens_before_target. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> | 23 天前 | |
chore(skill): keep __pycache__ out of deployed skill directories (#222) * fix(skill): keep __pycache__ out of deployed skill directories When book-to-skill is installed as an agent skill, the skill *is* the source tree: it is cloned or unpacked into a host skills root (e.g. `~/.config/opencode/skills/book-to-skill/`) and executed from there. Running any tool therefore writes `__pycache__/*.pyc` *inside the skill*, turning a 27-file install into 43 files of payload — noise in diffs, extra surface to scan, and bytecode shipped wherever the skill is published or shared. The guard has to live in the executable entry points, not at module import. `book_to_skill/utils.py` is a pure library module (imported by the package, the CLI and the tests); assigning `sys.dont_write_bytecode` at its import time silently mutates any embedding process that merely imports the package. So the policy is now wrapped in `if __name__ == "__main__":` — false for an embedding import — immediately after `import sys` and BEFORE `book_to_skill` is imported, in the four modules that are actually run directly: scripts/extract.py Step 2 extraction tools/scan_generated_skill.py Step 9.5 mandatory security scan tools/discovery_tax.py discovery-loop measurement book_to_skill/cli.py package CLI, runnable directly The import-time assignment in `book_to_skill/utils.py` is deleted outright. An agent invoking the skill does not set PYTHONDONTWRITEBYTECODE, so the guard has to live in the code rather than in the caller's environment. Two things that do NOT work, both measured: * Putting the guard in `book_to_skill/__init__.py` — Python compiles `__init__.py` before executing its body, so `__init__.cpython-*.pyc` is still emitted (27 -> 28 files). * `python -m book_to_skill` cannot be protected at all, for the same reason. It is not a documented path: SKILL.md invokes the extractor as `"$PYTHON_BIN" "$SCRIPT_PATH"`, never via `-m`. Adds tests/test_no_bytecode_pollution.py. It copies the skill payload into an isolated deployment (never the developer's tree), strips PYTHONDONTWRITEBYTECODE and PYTHONPYCACHEPREFIX, runs each entry point with `--help`, and asserts both that it exited 0 with its expected banner/usage on stdout/stderr and that no `.pyc`/`__pycache__` appeared. The old destructive `_clean()` helper, which deleted bytecode from the developer's checkout, is gone. A guard test diffs the discovered entry-point set against an explicit list, so a new invocable module fails the suite instead of silently shipping bytecode. A fourth test imports the package in a subprocess with `sys.dont_write_bytecode` pre-set to False and asserts it is unchanged, which is the direct regression guard for the import-time side effect removed here. Verified RED before the fix (entry points reporting "left build artifacts" and "no __main__ guard") and GREEN after, with a manual negative control: temporarily making an entry point raise SystemExit(1) fails the suite. * test(skill): prove the no-bytecode suite can fail The regression test could pass vacuously: it discarded the entry point's return code and asserted only that no bytecode files appeared, so an entry point that immediately raised SystemExit(1) still gave "1 passed". The old "the import really happened" probe ran in a *separate* subprocess, so it could not prove the invocation under test worked either. Add `test_negative_control_broken_entry_point_is_reported_as_failure`: it builds a deliberately broken entry point (body immediately `raise SystemExit(1)`, producing no output) and runs it through the *same* `_run_entry` helper the real tests use, asserting the helper reports failure. Because that helper checks both `returncode == 0` and the presence of the tool's expected output, this test locking in an `AssertionError` is the proof that "returncode is discarded" is no longer possible for any entry point. Demonstrated by temporarily breaking a real entry point: the suite fails with "scripts/extract.py exited 1; it must run successfully (--help)" and passes again once restored. * test(skill): avoid implicit string concatenation in import probe CodeQL flagged the list of implicitly concatenated string literals as a possible missing comma. Use explicit newlines in a single parenthesised string instead, which reads the same and removes the warning. --------- Co-authored-by: wgm66 <wgm66@users.noreply.github.com> | 17 天前 | |
chore(skill): keep __pycache__ out of deployed skill directories (#222) * fix(skill): keep __pycache__ out of deployed skill directories When book-to-skill is installed as an agent skill, the skill *is* the source tree: it is cloned or unpacked into a host skills root (e.g. `~/.config/opencode/skills/book-to-skill/`) and executed from there. Running any tool therefore writes `__pycache__/*.pyc` *inside the skill*, turning a 27-file install into 43 files of payload — noise in diffs, extra surface to scan, and bytecode shipped wherever the skill is published or shared. The guard has to live in the executable entry points, not at module import. `book_to_skill/utils.py` is a pure library module (imported by the package, the CLI and the tests); assigning `sys.dont_write_bytecode` at its import time silently mutates any embedding process that merely imports the package. So the policy is now wrapped in `if __name__ == "__main__":` — false for an embedding import — immediately after `import sys` and BEFORE `book_to_skill` is imported, in the four modules that are actually run directly: scripts/extract.py Step 2 extraction tools/scan_generated_skill.py Step 9.5 mandatory security scan tools/discovery_tax.py discovery-loop measurement book_to_skill/cli.py package CLI, runnable directly The import-time assignment in `book_to_skill/utils.py` is deleted outright. An agent invoking the skill does not set PYTHONDONTWRITEBYTECODE, so the guard has to live in the code rather than in the caller's environment. Two things that do NOT work, both measured: * Putting the guard in `book_to_skill/__init__.py` — Python compiles `__init__.py` before executing its body, so `__init__.cpython-*.pyc` is still emitted (27 -> 28 files). * `python -m book_to_skill` cannot be protected at all, for the same reason. It is not a documented path: SKILL.md invokes the extractor as `"$PYTHON_BIN" "$SCRIPT_PATH"`, never via `-m`. Adds tests/test_no_bytecode_pollution.py. It copies the skill payload into an isolated deployment (never the developer's tree), strips PYTHONDONTWRITEBYTECODE and PYTHONPYCACHEPREFIX, runs each entry point with `--help`, and asserts both that it exited 0 with its expected banner/usage on stdout/stderr and that no `.pyc`/`__pycache__` appeared. The old destructive `_clean()` helper, which deleted bytecode from the developer's checkout, is gone. A guard test diffs the discovered entry-point set against an explicit list, so a new invocable module fails the suite instead of silently shipping bytecode. A fourth test imports the package in a subprocess with `sys.dont_write_bytecode` pre-set to False and asserts it is unchanged, which is the direct regression guard for the import-time side effect removed here. Verified RED before the fix (entry points reporting "left build artifacts" and "no __main__ guard") and GREEN after, with a manual negative control: temporarily making an entry point raise SystemExit(1) fails the suite. * test(skill): prove the no-bytecode suite can fail The regression test could pass vacuously: it discarded the entry point's return code and asserted only that no bytecode files appeared, so an entry point that immediately raised SystemExit(1) still gave "1 passed". The old "the import really happened" probe ran in a *separate* subprocess, so it could not prove the invocation under test worked either. Add `test_negative_control_broken_entry_point_is_reported_as_failure`: it builds a deliberately broken entry point (body immediately `raise SystemExit(1)`, producing no output) and runs it through the *same* `_run_entry` helper the real tests use, asserting the helper reports failure. Because that helper checks both `returncode == 0` and the presence of the tool's expected output, this test locking in an `AssertionError` is the proof that "returncode is discarded" is no longer possible for any entry point. Demonstrated by temporarily breaking a real entry point: the suite fails with "scripts/extract.py exited 1; it must run successfully (--help)" and passes again once restored. * test(skill): avoid implicit string concatenation in import probe CodeQL flagged the list of implicitly concatenated string literals as a possible missing comma. Use explicit newlines in a single parenthesised string instead, which reads the same and removes the warning. --------- Co-authored-by: wgm66 <wgm66@users.noreply.github.com> | 17 天前 | |
feat(hosts): add Opencode skill discovery and validation (#216) * feat(hosts): add Opencode skill discovery and validation Adds OpenCode (sst/opencode) as a supported skill host, following the same shape as the Hermes Agent host: a validator lens, discovery paths in the converter spec, and tests for both. Validation (`tools/validate_skill.py --lens opencode`): - name must match the official pattern `^[a-z0-9]+(-[a-z0-9]+)*$` (lowercase alphanumeric with single hyphen separators) — stricter than the general Agent Skills default, so dots and underscores are rejected - only the five documented frontmatter keys are recognized: name, description, license, compatibility, metadata - no `allowed-tools` enforcement: OpenCode grants access through `permission.skill` in opencode.json, not frontmatter - no description soft limit; only the 1024-char hard cap applies Discovery (SKILL.md): - personal roots: ~/.config/opencode/skills (OpenCode-managed), plus the ~/.agents/skills and ~/.claude/skills compatibility roots - project roots: .opencode/skills, .agents/skills, .claude/skills - no trust gate, unlike Hermes: OpenCode discovers skills by walking up from the working directory to the git worktree Tests: - 11 `test_opencode_lens_*` cases in tests/test_validate_skill.py - tests/test_opencode_host_support.py, symmetric to the Hermes suite minus its trust-gate cases (OpenCode has no such model); guarded by a real bash probe so it skips instead of failing on hosts where the WSL stub makes `bash` unusable Docs: install.md, architecture.md, how-it-works.md, index.md and the three READMEs list OpenCode alongside the existing hosts. The README repository map also listed only three lenses; it now lists all five. Refs: https://opencode.ai/docs/skills * test(opencode): probe project layout from the project root The `project-opencode` case failed on CI (Linux) while passing locally, because the new host suite is skipped on hosts without a working bash — so the guard hid a real bug until CI ran it. The spec's extractor probe lists project-local roots other than Hermes' as CWD-relative paths (`.opencode/skills/...`, alongside `.github/`, `.claude/`, `.agents/`); only the Hermes block rewrites candidates against `$PROJECT_ROOT`. The test ran the project layout from `project/src/nested`, where no `.opencode/` exists, so `SCRIPT_PATH` stayed empty and the assertion compared the CWD against the extractor. Run project layouts from the project root and keep the personal layouts at a nested CWD, which is what actually demonstrates that personal roots resolve as absolute `$HOME/...` paths independent of depth. Verified under a real bash: the project case now resolves, and the three personal cases (`.config/opencode`, `.agents`, `.claude`) still resolve from `project/src/nested`. * fix(hosts): resolve project-local skills from the git worktree Project-local candidates were listed as bare CWD-relative paths (`.opencode/skills/...`, `.github/...`, `.claude/...`, `.agents/...`), so the converter was only found when the agent happened to sit at the project root. From a nested directory such as `project/src/nested` the probe returned nothing and Step 2 failed with "Could not find scripts/extract.py". That contradicts how the hosts behave: OpenCode walks up from the working directory to the git worktree, so project skills are meant to be reachable from anywhere inside the project. Prepend the same four roots resolved against `$PROJECT_ROOT` (already computed by the probe) when it is non-empty, keeping the CWD-relative forms afterwards so behaviour outside a git repository is unchanged: ```bash if [ -n "$PROJECT_ROOT" ]; then CANDIDATES+=( "$PROJECT_ROOT/.github/skills/book-to-skill/scripts/extract.py" "$PROJECT_ROOT/.claude/skills/book-to-skill/scripts/extract.py" "$PROJECT_ROOT/.agents/skills/book-to-skill/scripts/extract.py" "$PROJECT_ROOT/.opencode/skills/book-to-skill/scripts/extract.py" ) fi ``` Tests: the nested-directory case is now asserted rather than worked around. `test_opencode_extractor_probe_discovers_supported_layouts` runs every layout from `project/src/nested`, and a new `test_opencode_project_probe_resolves_project_root_from_nested_cwd` covers `.opencode`, `.agents` and `.claude` from `project/src/nested/deeper`. The previous version ran the project layout from the project root, which hid the bug instead of catching it. Also fixes two test-infrastructure faults that made these tests skip on Windows and therefore unable to catch anything: * Bash detection invoked a bare `bash`, which CreateProcess resolves against the system directory before PATH, always picking the WSL stub (`C:\Windows\System32\bash.exe`) even though a working Git bash exists. Now tries absolute candidates — `$OPENCODE_GIT_BASH_PATH`, Git's `../bin/bash.exe`, `/usr/bin/bash` — and validates with `bash -c true`. * Paths returned by Git bash are in MSYS form (`/tmp/...` for `$HOME`) and never compared equal to `pathlib` paths. `_resolve()` now normalises through `cygpath -w` on Windows. Negative control: reverting the SKILL.md change fails exactly the four nested/project cases (4 failed, 12 passed); restoring it returns 16 passed. * fix(hosts): resolve review findings on the rebased selection policy - SKILL.md: remove the duplicated pre-#125 rule-6 orphan left by the conflict resolution; add OpenCode to the host prompt (rule 6) so the unknown-host question names every supported host; repair the frontmatter delimiter that the rebase glued to the description line. - tests/test_scope_selection_contract.py: stop freezing the selection list at exactly six rules and the pre-#209 prompt sentence; assert one contiguous 1..N list plus the required policy instead (review request, 2026-09-22). - tests/test_opencode_host_support.py: same de-freezing for the unknown-host prompt assertion (#209 legitimately grew that list). - tools/validate_skill.py: fix the lens table so the OpenClaw entry keeps its own keys and OpenCode no longer carries a duplicated enforces_allowed_tools/name_pattern pair. Windows-local gates: 695 passed, 16 skipped; the 17 test_hermes failures are the documented WSL-stub bash artifact (Linux CI green). ruff clean; validate_skill passes for claude and opencode lenses. --------- Co-authored-by: wgm66 <wgm66@users.noreply.github.com> | 10 天前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 23 天前 | ||
| 17 天前 | ||
| 17 天前 | ||
| 10 天前 |