crawlee:端到端网络爬虫与数据抓取工具,模拟人类行为轻松绕过反爬机制

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

分支74Tags646

项目介绍

Crawlee —— 用于Node.js的网页抓取和浏览器自动化库,助您快速构建可靠的爬虫。【此简介由AI生成】

定制我的领域
13124.84 K1.57 K访问 GitHub