A command-line tool to crawl websites using puppeteer.
一个网络安全法律法规、安全政策、国家标准、行业标准知识库。A knowledge base of cybersecurity laws and regulations, security policies, national standard...
Crawl and convert any website into clean markdown
Parsed data from website https://jadwalsholat.org
Grabs current REWE discounts and saves them in a markdown file || Holt sich aktuelle REWE-Angebote und exportiert sie in eine Markdown-Liste
A universal solution for web crawling lists. 抓取网页列表的通用解决方案
Turn any developer documentation into a GPT
Enterprise-grade, multi-threaded SEO Crawler, Rule Engine, and Link Graph Analyzer. Built in TypeScript for speed, compliance, and deep site health au...
Java 網路資料爬蟲包
使用asyncio和aiohttp开发的轻量级异步协程web爬虫框架
Google Arts & Culture high quality image downloader
Continuously search imageboards threads for images/webms and download them
练手项目:Comment of Interest 电商文本评论数据挖掘 (爬虫 + 观点抽取 + 句子级和观点级情感分析)
A package to get list of user agents based on filters such as operating system, software name etc..
Python: An all-in-one Web Crawler, Web Parser and Web Scrapping library!
使 scrapy 开发不用在意 item,pipeline,middleware 等通用场景下模块的编写,解放开发者的双手。
A self-healing scraper for hostile sites: broken selectors repair themselves, browser rendering kicks in when needed, and a coherent identity layer (C...
爬取B站历史弹幕/全弹幕, 支持高级弹幕, Bas弹幕爬取. [2025年]可用; 内部爬取算法可以在 最优最少 请求次数下爬取弹幕, 并且 不会 丢失任何弹幕. 支持多任务管...
A spider on Dcard. Strong and speedy.
A multiprocessing crawler for weibo albums.
Discover hidden deepweb pages
👋 HOLA! ENJOY OUR GOOGLE MAPS SCRAPER 🚀 TO EFFORTLESSLY EXTRACT DATA SUCH AS NAMES, ADDRESSES, PHONE NUMBERS, WEBSITES, AND RATINGS FROM GOOGLE MAPS...
Scrapy-based Crawlers for news of Taiwan
🤖 A curated list of websites that restrict access to AI Agents, AI crawlers and GPTs
Calendar of Public Holidays in China 中国大陆节假日日历订阅 自动节假日闹钟
爬虫工程师常用的 Chrome 插件 | Chrome extensions used by crawler developer
Fast, highly configurable, cloud native dark web crawler.
:spider: This is an ES6 adaptation of the original PHP library CrawlerDetect, this library will help you detect bots/crawlers/spiders vie the useragen...
a scaleable and efficient crawelr with docker cluster , crawl million pages in 2 hours with a single machine
Fast, free, open-source technical SEO + GEO crawler: built for humans and agents.
Download audio tracks from Netflix to sample your favorite shows
lianjia / beike estate crawler/analysis 2024
An infinite Pinterest crawler/scraper. Crawl image with inifnite-scroll!
Crawl sites for RSS, Atom, and JSON feeds.
GOPA, a spider written in Go.(NOTE: this project moved to https://github.com/infinitbyte/gopa )
A collection of Python tools, scripts and utilities to make your life easier.
The LAW next generation crawler.
MediaCrawler is a powerful web scraper for self-media platforms. Easily collect and analyze content to enhance your digital strategy. 🌐🕷️
医疗知识图谱构建实战,通过爬虫获取百度百科数据,使用Mongodb存储结构化三元组,并使用neo4j进行知识图谱的构建及可视化; Medical Knowledge Graph; Crawler;...
All In One, Fast, Easy Recon Tool
Price tracker of Amazon
Selenium automation test framework
A simple Elixir library for writing decently-performing crawlers with minimum effort.
这是一个用Python写的小说爬虫软件
extract data from html table
使用 SpringBoot2.0+ElasticSearch 实现的开源电影搜索引擎
ScrapeGPT is a RAG-based Telegram bot designed to scrape and analyze websites, then answer questions based on the scraped content. The bot utilizes Re...
Extract structured data from Shopify websites.
Python-based web crawling script with randomized intervals, user-agent rotation, and proxy server IP rotation to outsmart website bots and prevent blo...
爬取及整理Freebuf\安全客\先知\知道创宇等站点的”web安全“类优质文章