Web Scraping & Crawling

Frameworks and tools for web scraping, crawling, and data extraction

Web Scraping & Crawling — comparison of firecrawl, Scrapling, crawl4ai, scrapy, crawlee
Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Scrapy, a fast high-level web crawling & scraping framework for Python.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Popularity
Stars190,11586,64585,13764,70926,100
Global Rank#35#183#185#324#1561
Weekly Activity(Oct 3 – Oct 9)
New Stars+2,420+1,344+508+161+129
Pushes1100512
Issues Closed00000
Community
Forks10,0838,8788,83411,9961,698
Contributors1753295749152
Open Issues53814237265147
Project Info
OwnerfirecrawlD4Vinciunclecodescrapyapify
LicenseAGPL-3.0BSD-3-ClauseApache-2.0BSD-3-ClauseApache-2.0
LanguageTypeScriptPythonPythonPythonTypeScript
CreatedApr 2024Oct 2024May 2024Feb 2010Aug 2016