Internet-in-a-Box - Build your own LIBRARY OF ALEXANDRIA with a Raspberry Pi !
List of anti-detect and humanizing tools and browsers, including captcha solvers and sms-activation.
Anti-Detect Browser that passes every bot detection test. Drop-in Playwright replacement.
Hide your scrapers IP behind the cloud. Provision proxy servers across different cloud providers to improve your scraping success.
Scrape tweets, profiles, followers and following from Twitter/X, no API key needed. Python library with smart multi-account pooling, proxy support and...
Get clean data from tricky documents, powered by vision-language models ⚡
In this tutorial, we showcase how to scrape public Google data with Python and Oxylabs API.
A Modern Search Engine API for Anime, Movies/TVShows, Books, Light Novels, Manga, etc.
AgentQL is a suite of tools for connecting your AI to the web. Featuring a query language and Playwright integrations for interacting with elements an...
Example end to end data engineering project.
An ergonomic, privacy-aware Python HTTP Client
Collection of patches for puppeteer and playwright to avoid automation detection and leaks. Helps to avoid Cloudflare and DataDome CAPTCHA pages. Easy...
🤖 Scrape data from HTML websites automatically by just providing examples
📰 Diários oficiais brasileiros acessíveis a todos | 📰 Brazilian government gazettes, accessible to everyone.
Parsel lets you extract data from XML/HTML documents using XPath or CSS selectors
Open Source Bulk Auto Gmail Creator Bot with Selenium & Seleniumwire ( Python ). Feel free to contact me with Django/Flask, ML, AI, GPT, Automation, S...
Lightweight library for scraping web-sites with LLMs
Watch everything from your terminal.
This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster.
HTTP(S)/SOCKS5 rotating residential proxies - code examples & general information.
🎭 Intelligent browser header & fingerprint generator
Tools for various online judges. Downloading sample cases, generating additional test cases, testing your code, and submitting it.
:rocket: An open source alternative to searx which provides a modern-looking :sparkles:, lightning-fast :zap:, privacy respecting :disguised_face:, se...
Creating Scrapy scrapers via the Django admin interface
📰 Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.
A completely revamped and redesigned fork, reimagined from scratch based on the original onlyfans-scraper
artoo.js - the client-side scraping companion.
Crawly, a high-level web crawling & scraping framework for Elixir.
modular service framework to move and transform network packets
Automatically archive links to videos, images, and social media content from Google Sheets (and more).
Your browser anime experience from the terminal
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Scalable Python web scraping scripts for +40 popular domains
🧹 Python package for text cleaning
🚀 Web scraping for humans
Scrape the Instagram frontend. Inspired from twitter-scraper by @kennethreitz.
The agent that turns websites into APIs!
Generate Free Edu Mail(s) within minutes
📄 Python tool to turn Notion.so pages into lightweight, customizable static websites
A CLI toolset to generate table of contents for PDF files automatically.
Simple but useful Python web scraping tutorial code.
DataHen Till is a companion tool to your existing web scraper that instantly makes it scalable, maintainable, and more unblockable, with minimal code...
Kuwala is the no-code data platform for BI analysts and engineers enabling you to build powerful analytics workflows. We are set out to bring state-of...
[Unmaintained] A simple and clean video/music/image downloader 👾
:scissors: High performance, multi-threaded image scraper
Lookyloo is a web interface that allows users to capture a website page and then display a tree of domains that call each other.
🕵️♂️ LinkedIn profile scraper returning structured profile data in JSON.
🥫 The simple, fast, and modern web scraping library
😈📚 A curated library of research papers and presentations for counter-detection and web privacy enthusiasts.
Google Search Results via SERP API pip Python Package