| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat(extraction): gate Reddit + Amazon site_data behind anti-bot block detection Audit C5: Reddit and Amazon URLs that returned anti-bot challenges or "Page Not Found" landings were silently treated as successful empty site_data responses — callers had no way to tell the difference between a real empty payload and a blocked page. - detectAntiBotBlock() exported from reddit + amazon site extractors, scans the first 10KB for canonical block / not-found phrases ("blocked by network security", "Page Not Found", robot-check prompts, etc.). - extract() short-circuits to null on a blocked body, refusing to emit fake site_data. - routedExtract surfaces the block reason via ExtractionResult.site_data_blocked. - handleFetch promotes it to FetchOutput.fetch_failed="blocked" so callers branch honestly while the fallback markdown body still ships. - Positive fixtures unchanged: real product / thread bodies still emit site_data with no false-positive blocked envelopes. | 4 个月前 | |
feat(extraction): add Amazon product site extractor (C3) New site extractor under src/extraction/site-extractors/amazon.ts covers product pages on amazon.com, amazon.co.uk, amazon.de (plus all major ccTLDs) and the amzn.to / a.co short-link domains. Returns the typed AmazonProduct shape (asin, title, brand, price, currency, rating, review_count, description, features, specifications, images, availability) via extractAmazonProduct(), and adapts that into the existing ExtractionResult contract for the site-extractor registry. ASIN is derived from the URL path first (/dp/<asin>/, /gp/product/, etc.) and falls back to the data-asin DOM attribute when the URL is opaque (search-result snapshots, mobile renders). Price parsing prefers the .a-offscreen screen-reader text — Amazon's canonical formatted price — and falls back to the visible whole/fraction split when offscreen is missing. Currency symbols are mapped to ISO 4217 via a small lookup table so callers do not have to handle localisation. Out-of-stock pages return price=null (never 0) so price-based sorting never falsely ranks them as cheapest. Image filtering drops data: URIs and tracking pixels. Fixtures (each <100KB) cover electronics, books, groceries, an out-of-stock product, and a GBP-priced UK product. 45 new tests; every test name encodes the why so the suite is regression-catchable. | 4 个月前 | |
fix(c3): proto-pollution guard on specifications + EUR/x-locale/spoofed-host tests - specifications: use Object.create(null) and skip __proto__/constructor/prototype keys in parseOverviewSpecifications + parseDetailBulletSpecifications so an attacker-controlled label cannot smuggle reserved slots into the result - parseFeatures: soft-cap at 100 entries to harden against adversarial bullet floods - tests: spoofed amazon.com.attacker.com host rejection, EUR fixture (de-euro.html), /x-locale/ image filter, proto-pollution guards on both spec sources + global Object.prototype probe, features soft-cap probe | 4 个月前 | |
feat(extraction): add Amazon product site extractor (C3) New site extractor under src/extraction/site-extractors/amazon.ts covers product pages on amazon.com, amazon.co.uk, amazon.de (plus all major ccTLDs) and the amzn.to / a.co short-link domains. Returns the typed AmazonProduct shape (asin, title, brand, price, currency, rating, review_count, description, features, specifications, images, availability) via extractAmazonProduct(), and adapts that into the existing ExtractionResult contract for the site-extractor registry. ASIN is derived from the URL path first (/dp/<asin>/, /gp/product/, etc.) and falls back to the data-asin DOM attribute when the URL is opaque (search-result snapshots, mobile renders). Price parsing prefers the .a-offscreen screen-reader text — Amazon's canonical formatted price — and falls back to the visible whole/fraction split when offscreen is missing. Currency symbols are mapped to ISO 4217 via a small lookup table so callers do not have to handle localisation. Out-of-stock pages return price=null (never 0) so price-based sorting never falsely ranks them as cheapest. Image filtering drops data: URIs and tracking pixels. Fixtures (each <100KB) cover electronics, books, groceries, an out-of-stock product, and a GBP-priced UK product. 45 new tests; every test name encodes the why so the suite is regression-catchable. | 4 个月前 | |
feat(extraction): add Amazon product site extractor (C3) New site extractor under src/extraction/site-extractors/amazon.ts covers product pages on amazon.com, amazon.co.uk, amazon.de (plus all major ccTLDs) and the amzn.to / a.co short-link domains. Returns the typed AmazonProduct shape (asin, title, brand, price, currency, rating, review_count, description, features, specifications, images, availability) via extractAmazonProduct(), and adapts that into the existing ExtractionResult contract for the site-extractor registry. ASIN is derived from the URL path first (/dp/<asin>/, /gp/product/, etc.) and falls back to the data-asin DOM attribute when the URL is opaque (search-result snapshots, mobile renders). Price parsing prefers the .a-offscreen screen-reader text — Amazon's canonical formatted price — and falls back to the visible whole/fraction split when offscreen is missing. Currency symbols are mapped to ISO 4217 via a small lookup table so callers do not have to handle localisation. Out-of-stock pages return price=null (never 0) so price-based sorting never falsely ranks them as cheapest. Image filtering drops data: URIs and tracking pixels. Fixtures (each <100KB) cover electronics, books, groceries, an out-of-stock product, and a GBP-priced UK product. 45 new tests; every test name encodes the why so the suite is regression-catchable. | 4 个月前 | |
feat(extraction): add Amazon product site extractor (C3) New site extractor under src/extraction/site-extractors/amazon.ts covers product pages on amazon.com, amazon.co.uk, amazon.de (plus all major ccTLDs) and the amzn.to / a.co short-link domains. Returns the typed AmazonProduct shape (asin, title, brand, price, currency, rating, review_count, description, features, specifications, images, availability) via extractAmazonProduct(), and adapts that into the existing ExtractionResult contract for the site-extractor registry. ASIN is derived from the URL path first (/dp/<asin>/, /gp/product/, etc.) and falls back to the data-asin DOM attribute when the URL is opaque (search-result snapshots, mobile renders). Price parsing prefers the .a-offscreen screen-reader text — Amazon's canonical formatted price — and falls back to the visible whole/fraction split when offscreen is missing. Currency symbols are mapped to ISO 4217 via a small lookup table so callers do not have to handle localisation. Out-of-stock pages return price=null (never 0) so price-based sorting never falsely ranks them as cheapest. Image filtering drops data: URIs and tracking pixels. Fixtures (each <100KB) cover electronics, books, groceries, an out-of-stock product, and a GBP-priced UK product. 45 new tests; every test name encodes the why so the suite is regression-catchable. | 4 个月前 | |
feat(extraction): add Amazon product site extractor (C3) New site extractor under src/extraction/site-extractors/amazon.ts covers product pages on amazon.com, amazon.co.uk, amazon.de (plus all major ccTLDs) and the amzn.to / a.co short-link domains. Returns the typed AmazonProduct shape (asin, title, brand, price, currency, rating, review_count, description, features, specifications, images, availability) via extractAmazonProduct(), and adapts that into the existing ExtractionResult contract for the site-extractor registry. ASIN is derived from the URL path first (/dp/<asin>/, /gp/product/, etc.) and falls back to the data-asin DOM attribute when the URL is opaque (search-result snapshots, mobile renders). Price parsing prefers the .a-offscreen screen-reader text — Amazon's canonical formatted price — and falls back to the visible whole/fraction split when offscreen is missing. Currency symbols are mapped to ISO 4217 via a small lookup table so callers do not have to handle localisation. Out-of-stock pages return price=null (never 0) so price-based sorting never falsely ranks them as cheapest. Image filtering drops data: URIs and tracking pixels. Fixtures (each <100KB) cover electronics, books, groceries, an out-of-stock product, and a GBP-priced UK product. 45 new tests; every test name encodes the why so the suite is regression-catchable. | 4 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 4 个月前 | ||
| 4 个月前 | ||
| 4 个月前 | ||
| 4 个月前 | ||
| 4 个月前 | ||
| 4 个月前 | ||
| 4 个月前 |