The context API to search, scrape, and interact with the web at scale. 🔥
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Scrapy, a fast high-level web crawling & scraping framework for Python.
🕵️♂️ Collect a dossier on a person by username from 3000+ sites
Open source AI job application bot in Python: auto apply to jobs, with a tailored resume and cover letter for each posting.
Python scraper based on AI
Elegant Scraper and Crawler Framework for Golang
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs,...
Pythonic HTML Parsing for Humans™
Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)
A scalable web crawler framework for Java.
🦊 Anti-detect browser
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, P...
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
List of libraries, tools and APIs for web scraping and data processing.
A Smart, Automatic, Fast and Lightweight Web Scraper for Python
Tabula is a tool for liberating data tables trapped inside PDF files
Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
🚀 Free HTTP, SOCKS4, & SOCKS5 proxy list * Updated every 5 minutes * and rotating proxy API (100+ countries)
Declarative data automation language and Go runtime for structured extraction workflows.
Convert cURL commands to Python, JavaScript, Java, C#, PHP, Go, Dart, R, Ruby, Rust, MATLAB, Elixir, CFML, Ansible, Strest or JSON
Distributed crawler powered by Headless Chrome
Self-hosted webscraper.
Swiss-army tool for scraping and extracting data from online assets, made for hackers
Twitter API Scraper | Without an API key | Twitter Internal API | Free | Twitter scraper | Twitter Bot
Mechanize is a ruby library that makes automated web interaction easy.
Collection of useful data science topics along with articles, videos, and code
Up-to-date simple useragent faker with real world database
Snoop — инструмент разведки на основе открытых данных (OSINT world)
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Nat...
Do you want to LEARN NEW STUFF for FREE? Don't worry, with the power of web-scraping and automation, this script will find the necessary Udemy coupons...
Scrape Facebook public pages without an API key
A browser testing and web crawling library for PHP and Symfony
An open source fingerprint browser based on Ungoogled Chromium. 指纹浏览器 隐私浏览器
A Python module to scrape several search engines (like Google, Yandex, Bing, Duckduckgo, ...). Including asynchronous networking support.
Learn step-by-step how to scrape Google Trends data and make a result comparison using Python and Oxylabs SERP API. Extract keywords, their popularity...
Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering.
Agentic Email Automation Tool: Describe your product. Define your target market. The AI finds the leads for you.
Python library and CLI for X/Twitter scraping with multi-account rotation and built-in rate-limit handling.
Get web data for AI agents and LLMs
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Advanced Privacy Browser Core with Unified Fingerprint Defense: Cloudflare, Akamai, Kasada, Shape, DataDome, PerimeterX, hCaptcha, FunCaptcha, Imperva...
A curated list of awesome puppeteer resources.
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
A CLI utility for taking screenshots of websites, recording video demos and scraping sites using JavaScript
Web Scraping Framework
Getting started with Puppeteer and Chrome Headless for Web Scraping
Get info from any web service or page
OSINT cheat sheet, list OSINT tools, wiki, dataset, article, book , red team OSINT for hackers and OSINT tips and OSINT branch. This repository will g...