| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat: Defuddle and Readability extraction wrappers | 5 个月前 | |
test: add HTML fixtures and integration tests for Trafilatura pipeline | 5 个月前 | |
feat: Defuddle and Readability extraction wrappers | 5 个月前 | |
test(extraction): lock the sparsity gate against dense ordinal leaderboards The committed negative fixtures all lacked rank-ordinal rows, so nothing exercised the sparsity gate (filled/gridCells < 0.5) — the last line of defense that stops a dense ordinal leaderboard from over-firing the per-story segmentation. Behavior was correct, but a future relaxation of the 0.5 threshold would silently over-fire with zero failing test. Add a fixture-based negative: a header-less dense leaderboard where most rows are rank-led but two are non-ordinal, so the record-start rows are a minority (the starts<bodyRows and >=3-records checks both pass) and every cell is populated (filled/gridCells ≈ 1.0). The sparsity gate is the ONLY gate left to reject segmentation, so the test isolates it. Mutation-verified: loosening the threshold to <= 1.0 folds the leaderboard to rank/title/meta and fails the test. | 2 个月前 | |
fix(extraction): restore per-story segmentation for legacy nested-table listings Legacy nested-table listings (HN-class front pages) lay out each story as a CYCLE of rows — a rank+title row, a points/comments meta row, then a spacer — inside one header-less <table>. extractTables synthesised col_1..col_N headers and emitted every physical <tr> as its own row, so a listing became an interleaved run-on dump ({col_1:"1.",col_3:"Title"} then {col_2:"342 points…"} then an all-empty spacer). That dump is useless to an agent AND, carrying only col_N headers, matches no schema field so schema mode returns {} on the page. The multi-column interleaved case was never handled: the round-1 degenerate salvage only fired for single-column (headers.length<=1) run-on tables, and the round-2 div-grid detector explicitly excludes <table>/<tr>/<td> tags, so a multi-<tr>-per-record listing fell straight through to the raw col_N mapping. Detect the pattern at the extractTables seam and collapse it to ONE row per record. The gate is per-structure and narrow: header-less table (no <th>) + >=3 rank-ordinal record starts that are a strict minority of rows + over half the row×column grid empty (sparsity). A dense data grid, a sparse non-ordinal layout, and a <2-record table all pass through untouched. Real legacy tables (DistroWatch rankings, FluxBB topic list, OpenBSD errata, the RFC index) do not fire. Schema mode returns clean records as a consequence, no schema.ts change. The div-grid pricing path (divs, not tables) is structurally disjoint and stays green. Both behaviours are locked in a permanent both-ways regression test with non-HN fixtures so a future round cannot silently trade one for the other. | 2 个月前 | |
test: add integration tests and job listing fixture for schema extraction | 5 个月前 | |
feat: add JSON-LD parser with @graph support and schema matching | 5 个月前 | |
fix(extraction): restore per-story segmentation for legacy nested-table listings Legacy nested-table listings (HN-class front pages) lay out each story as a CYCLE of rows — a rank+title row, a points/comments meta row, then a spacer — inside one header-less <table>. extractTables synthesised col_1..col_N headers and emitted every physical <tr> as its own row, so a listing became an interleaved run-on dump ({col_1:"1.",col_3:"Title"} then {col_2:"342 points…"} then an all-empty spacer). That dump is useless to an agent AND, carrying only col_N headers, matches no schema field so schema mode returns {} on the page. The multi-column interleaved case was never handled: the round-1 degenerate salvage only fired for single-column (headers.length<=1) run-on tables, and the round-2 div-grid detector explicitly excludes <table>/<tr>/<td> tags, so a multi-<tr>-per-record listing fell straight through to the raw col_N mapping. Detect the pattern at the extractTables seam and collapse it to ONE row per record. The gate is per-structure and narrow: header-less table (no <th>) + >=3 rank-ordinal record starts that are a strict minority of rows + over half the row×column grid empty (sparsity). A dense data grid, a sparse non-ordinal layout, and a <2-record table all pass through untouched. Real legacy tables (DistroWatch rankings, FluxBB topic list, OpenBSD errata, the RFC index) do not fire. Schema mode returns clean records as a consequence, no schema.ts change. The div-grid pricing path (divs, not tables) is structurally disjoint and stays green. Both behaviours are locked in a permanent both-ways regression test with non-HN fixtures so a future round cannot silently trade one for the other. | 2 个月前 | |
test(extract): integration tests for extract pipeline | 5 个月前 | |
feat: Defuddle and Readability extraction wrappers | 5 个月前 | |
test: add HTML fixtures and integration tests for Trafilatura pipeline | 5 个月前 | |
feat: add JSON-LD parser with @graph support and schema matching | 5 个月前 | |
test(extract,crawl): integration fixtures for markdown post-process | 4 个月前 | |
test(extraction): fail-first for react.dev reference nav-only regression Real served react.dev/reference/react HTML reproduces the nav-only failure: the reference body (intro prose + Hooks/Components/APIs index) is dropped and only the 7-link top nav survives (~228-char markdown). Adds the real-HTML fixture plus an integration test (full V1 provider boundary) and a boilerplate unit test pinning that a content-grid wrapper whose Tailwind class merely contains the substring 'sidebar' (grid-cols-sidebar-content) must not be deleted, while genuine sidebars still are. Both fail on current code. | 3 个月前 | |
test(extraction): SPA reference page returns body content at small cap | 3 个月前 | |
feat: Defuddle and Readability extraction wrappers | 5 个月前 | |
fix(extraction): guard main-containing wrappers in boilerplate strip VitePress nests the doc body as <div class="VPContent has-sidebar"> > <div class="VPContentDoc has-aside has-sidebar"><main>...</main></div>. The boilerplate pre-pass selector matched the has-sidebar state class on those layout wrappers and removed the whole content region before content-root isolation ran, leaving only the VitePress navbar cluster (~nav-only markdown) for vuejs.org/guide/introduction. Add a semantic guard to stripBoilerplateDom: never remove an element that contains the page's single <main> landmark. Boilerplate (nav/sidebar/ footer/feedback) sits beside <main>, never wraps it. The guard subsumes react.dev's grid-cols-sidebar-content case, so the sidebar selector is simplified back to plain [class*="sidebar"]. Guard keys on <main> only (not <article>) so related-content asides wrapping <article> cards stay removable. Empirical (real served HTML): vuejs guide body 220 -> 10535 chars (Single-File/declarative/Progressive present); react.dev reference body unchanged at 3611 chars. | 3 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 2 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 5 个月前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 3 个月前 | ||
| 5 个月前 | ||
| 3 个月前 |