Dynamic configurable crawl (动态可配置化爬虫)
web scraping extension
A Python crawler tool that can automatically simulate browser operations to crawl all users' tweet content and save all static resources (videos, pict...
Crawl & Visualize ICLR 2023 Data from OpenReview
An intelligent web service to automatically detect web content and extract information from it.
带图形界面的哔哩哔哩漫画下载工具
一个用于 哔哩哔哩漫画 B漫 的下载脚本
Crawler behind the Shopify App Marketplace dataset
Verify that a request is from Google crawlers using Google's DNS verification steps
Google Maps crawler using Selenium. All extracted data is forwarded to a SQS queue.
Fast website scraper and wordlist generator
NEU new ipgw python manager
xSMTP 🦟 Lightning fast, multithreaded smtp scanner targeting open-relay and unsecured servers in multiple network ranges.
This repo is a part of blog series on several web scraping projects where we will explore scraping techniques to crawl data from simple websites to we...
PHP Metacritic API - Mirror from my GitLab
Crawlzone is a fast asynchronous internet crawling framework for PHP.
JS逆向技巧:页面跳转JS代码定位通杀方案
Backup your friends' Instagram Stories forever and get to keep them even after 24 hours.
Golang爬虫 爬取汽车之家 二手车产品库
深圳大学抢课系统
Search through all your personal data efficiently like web search.
ClaudeChrome - Native browser context awareness for agents.
超高速异步协程Python爬虫
Parse government documents into well formed JSON
百度贴吧分布式爬虫,用于贴吧数据挖掘。从贴吧维度和用户维度进行数据分析
Powerful web scraping framework for Crystal
基于NodeJS的基金数据爬虫,爬取的数据存于github的@nullpointer/fund-data。
雪球爬虫 高效爬取近期沪深A股股票评论并自动生成PDF版情感分析报告
a puppeteer walker 🕷 🕸
Scrape Learning (ctrip)
Some scrapy and web.py exmaples
Slideshare to PDF downloader. Using Selenium and auto scroll-down to get the entire slides completely.
将任意网站转换为 RSS 订阅源、多引擎抓取、反爬突破、支持爬取抖音、快手、小红书、B站、知乎、小宇宙、知识星球
NTU CEIBA 資料下載工具
fetchman is a simple crawler system/简单好用的爬虫框架
tumblr解析网站
Automatically download all PDF files of searching results & their patent families found on Google Patents.
🔥 PHP library to warm up caches of URLs located in XML sitemaps
台股機器人,提供即時個股及大盤報價、走勢、新聞、盤後資料等 Telegram bot to query real-time TW stock quotes, charts, news, and other related informatio...
A PasteBin scrapper that doesnt rely on the PasteBin scrape API
Just a crawler based on tg-cli for Telegram. Deprecated by now, please use telegram-export.
python requests + Django+ nodejs koa+ mysql to crawl eastmoney fund and stock data,for data analysis and visualiaztion .
Vietnamese text data crawler scripts for various sites (including Youtube, Facebook, 4rum, news, ...)
Your preferred open source focused crawler for the deep web.
steam 特惠游戏榜单(自动刷新)
A collection of browser fingerprinting projects, research, and resources. Intended as a way to aggregate research surrounding the subject.
Midnight Sea: navigating in the waters of dark web markets
Privacy Web Search Engine (not meta, own crawler)
Extract data from websites in LLM ready JSON or CSV format. Crawl or Scrape entire website with Website Crawler
A bot that automatically sends emails to new ads posted in any desired xe.gr search url.