Web crawler and scraper based on Scrapy and Playwright's headless browser.
MCP server that enables self-healing automatic repair of Scrapy spiders. When websites change, your scrapers fix themselves.
Desktop app that crawls urls from Google's search engine results
Python module allowing you to do various searches for links on the Web.
web crawling & scraping framework for Python
Node.js tool for downloading all free MIDI files on VGMusic.com
Replayable Browser Agent
A basic tutorial to web scraping using python for beginners
converts webpage content into Markdown format, optimized for LLM training and context
A simple web crawller in go
Генератор сырых дампов пользователей VK.
crawling facebok page
Fetch, store and access user agent strings for different browsers
re-employment-kraken scrapes (job) sites, remembers what it saw and notifies downstream systems of any new sightings.
2023.11) velog statistics dashboard fullstack
Crawl and track followers count of Twitter account
A lightweight frontend for self-hosted Firecrawl instances
Engine for collecting onion domains and crawling from webpage based on Tor network
An interactive Command-Line Interface Build in NodeJS for downloading a single or multiple images to disk from URL
The definitive list of the latest libraries, tools, APIs and providers for web scraping. The only daily-updated collection of web scraping resources.
An ultra small PoC to show how to combine Apache Nutch and Apache Solr, crawling through web pages and storing the results in Solr for quering
Fast, parallel and easy to use web crawler for penetration testing and bug bounty
Crawling route waypoints for HK bus routes
App to scrap the web, for people without coding skills. Fully integrates WebCrawlers (Headless Chrome) and the interface to deal with it.
🚀 OMKAR TEMP MAIL HELPS YOU USE TEMPORARY EMAILS. 🤖
Đồ án cuối kì môn khoa học dữ liệu ứng dụng. Thu thập data bằng cách parsing HTML và sử dụng các mô hình học máy để giải quyết câu hỏi được đặt ra ba...
Intelligent web discovery agent with LLM-powered planning, multi-source search, smart deduplication, and GRPO preference dataset collection. Autonomou...
A full stack application that scrapes & filters YouTube comments using Google's Puppeteer, instead of using the YouTube API
Telegram channel relations analyzer
Fast extraction of all external links from wikipedia
东方财富网股票数据爬取
A sophisticated data-driven system that revolutionizes product discovery for dropshipping businesses. Unlike traditional web crawlers, this platform l...
A django application for scraping properties with scrapy.
Simple Manga Downloader, a tool to search and download manga
You Can Download Instagram Post With This Script
This is a crawler for crawling papers from google scholar (http://scholar.google.com). Credits for this code goes to (https://github.com/ckreibich/sch...
This article covers everything you need to get started with Crawlee. Learn more about its benefits and see a working example of scraping a website wit...
Python scripts, first traverses chrome Bookmark file and second removes stale entries. Includes Jenkinsfile to generate docker images.
[ACL 2024] Evaluation of the Fundus News Scraper
This Python script extracts comprehensive movie data from IMDB, focusing on top-grossing movies from 1920 to 2025. The scraper collects detailed infor...
🕷️ Easily scrap the web for torrent and media files.
Extraction, versioning and machine-readable provisioning of public data.
Crawler written in TypeScript using ES6 generators.
Simple scripts for crawling shopee's shop and product information from shopee.vn
Uses Sankey Diagrams to visualize politicians that have "crossed the floor" from election to election.
sᴇᴀʀᴄʜ ᴇɴɢɪɴᴇ sᴄʀᴀᴘᴇʀ ᴛᴏᴏʟ (ʙᴀsʜ)
crawling china stock recommendation from Sina Weibo, create pyecharts for data
Docker🐳 setup for automated news article crawling from German news websites. Written in Python🐍, uses MongoDB
Demonstration for crawling Laptop products on Tiki ecomercial website
Extract, crawl, and scrape web data efficiently with a powerful open-source CLI tool requiring no API keys and minimal setup.