diff --git a/collections/web-scraping/index.md b/collections/web-scraping/index.md new file mode 100644 index 00000000000..56385bccfb7 --- /dev/null +++ b/collections/web-scraping/index.md @@ -0,0 +1,30 @@ +--- +items: + - scrapy/scrapy + - apify/crawlee + - apify/crawlee-python + - gocolly/colly + - D4Vinci/Scrapling + - spider-rs/spider + - projectdiscovery/katana + - firecrawl/firecrawl + - unclecode/crawl4ai + - ScrapeGraphAI/Scrapegraph-ai + - microsoft/playwright + - puppeteer/puppeteer + - SeleniumHQ/selenium + - seleniumbase/SeleniumBase + - lightpanda-io/browser + - browserless/browserless + - cheeriojs/cheerio + - jhy/jsoup + - sparklemotion/nokogiri + - PuerkitoBio/goquery + - symfony/panther + - dgtlmoon/changedetection.io + - getmaxun/maxun + - crawlee-cloud/crawlee-cloud +display_name: Web scraping +created_by: aminembarki +--- +The web is the world's largest dataset, and these open source projects help developers collect it: crawling frameworks for Python, JavaScript, Go, Rust, Java, Ruby, and PHP, headless browser automation, HTML parsers, and self-hosted platforms for running crawlers at scale.