Topic

crawler

Repositories (1456)

pappet
pappet patrickschur JavaScript

A command-line tool to crawl websites using puppeteer.

105
awesome-chinese-law
awesome-chinese-law XiaomingX

一个网络安全法律法规、安全政策、国家标准、行业标准知识库。A knowledge base of cybersecurity laws and regulations, security policies, national standard...

105
firecrawl-py
firecrawl-py firecrawl Python

Crawl and convert any website into clean markdown

105
jadwalsholatorg
jadwalsholatorg lakuapik Python

Parsed data from website https://jadwalsholat.org

105
rewe-discounts
rewe-discounts foo-git Python

Grabs current REWE discounts and saves them in a markdown file || Holt sich aktuelle REWE-Angebote und exportiert sie in eine Markdown-Liste

104
crawlist
crawlist WwwwwyDev Python

A universal solution for web crawling lists. 抓取网页列表的通用解决方案

104
devdocs-to-llm
devdocs-to-llm alexfazio Jupyter Notebook

Turn any developer documentation into a GPT

104
SEOCORE
SEOCORE codepurse TypeScript

Enterprise-grade, multi-threaded SEO Crawler, Rule Engine, and Link Graph Analyzer. Built in TypeScript for speed, compliance, and deep site health au...

104
CrawlerPack
CrawlerPack abola Java

Java 網路資料爬蟲包

104
asyncpy
asyncpy lixi5338619 Python

使用asyncio和aiohttp开发的轻量级异步协程web爬虫框架

103
google-arts-crawler
google-arts-crawler piotrantosz Python

Google Arts & Culture high quality image downloader

103
4scanner
4scanner pboardman Python

Continuously search imageboards threads for images/webms and download them

103
COI
COI AlvinAi96 Jupyter Notebook

练手项目:Comment of Interest 电商文本评论数据挖掘 (爬虫 + 观点抽取 + 句子级和观点级情感分析)

102
random_user_agent
random_user_agent Luqman-Ud-Din Python

A package to get list of user agents based on filters such as operating system, software name etc..

102
webb
webb hardikvasa Python

Python: An all-in-one Web Crawler, Web Parser and Web Scrapping library!

102
AyugeSpiderTools
AyugeSpiderTools shengchenyang Python

使 scrapy 开发不用在意 item,pipeline,middleware 等通用场景下模块的编写,解放开发者的双手。

101
anansi
anansi mdowis Python

A self-healing scraper for hostile sites: broken selectors repair themselves, browser rendering kicks in when needed, and a coherent identity layer (C...

101
BiLiBiLi_DanMu_Crawling
BiLiBiLi_DanMu_Crawling HengXin666 TypeScript

爬取B站历史弹幕/全弹幕, 支持高级弹幕, Bas弹幕爬取. [2025年]可用; 内部爬取算法可以在 最优最少 请求次数下爬取弹幕, 并且 不会 丢失任何弹幕. 支持多任务管...

100
dcard-spider
dcard-spider leVirve Python

A spider on Dcard. Strong and speedy.

99
Weibo-Album-Crawler
Weibo-Album-Crawler Lodour Python

A multiprocessing crawler for weibo albums.

99
deepweb-scappering
deepweb-scappering kurogai Python

Discover hidden deepweb pages

98
google-maps-scraper
google-maps-scraper omkarcloud Python

👋 HOLA! ENJOY OUR GOOGLE MAPS SCRAPER 🚀 TO EFFORTLESSLY EXTRACT DATA SUCH AS NAMES, ADDRESSES, PHONE NUMBERS, WEBSITES, AND RATINGS FROM GOOGLE MAPS...

98
Taiwan-news-crawlers
Taiwan-news-crawlers TaiwanStat Python

Scrapy-based Crawlers for news of Taiwan

97
the-great-gpt-firewall
the-great-gpt-firewall samber Python

🤖 A curated list of websites that restrict access to AI Agents, AI crawlers and GPTs

97
chinese-holidays-calendar
chinese-holidays-calendar muhac Haskell

Calendar of Public Holidays in China 中国大陆节假日日历订阅 自动节假日闹钟

97
crawler-chrome-extensions
crawler-chrome-extensions zkqiang

爬虫工程师常用的 Chrome 插件 | Chrome extensions used by crawler developer

97
bathyscaphe
bathyscaphe creekorful Go

Fast, highly configurable, cloud native dark web crawler.

96
es6-crawler-detect
es6-crawler-detect JefferyHus TypeScript

:spider: This is an ES6 adaptation of the original PHP library CrawlerDetect, this library will help you detect bots/crawlers/spiders vie the useragen...

96
scaleable-crawler-with-docker-cluster
scaleable-crawler-with-docker-cluster tonywangcn Python

a scaleable and efficient crawelr with docker cluster , crawl million pages in 2 hours with a single machine

95
crawlie
crawlie spronta Rust

Fast, free, open-source technical SEO + GEO crawler: built for humans and agents.

95
narr
narr IljaN Go

Download audio tracks from Netflix to sample your favorite shows

95
lianjia-eroom-analysis
lianjia-eroom-analysis linpingta Python

lianjia / beike estate crawler/analysis 2024

95
Pinterest-infinite-crawler
Pinterest-infinite-crawler mirusu400 Python

An infinite Pinterest crawler/scraper. Crawl image with inifnite-scroll!

95
feedsearch-crawler
feedsearch-crawler DBeath Python

Crawl sites for RSS, Atom, and JSON feeds.

95
gopa-abandoned
gopa-abandoned medcl Go

GOPA, a spider written in Go.(NOTE: this project moved to https://github.com/infinitbyte/gopa )

94
python-tools
python-tools lucasayres Python

A collection of Python tools, scripts and utilities to make your life easier.

92
BUbiNG
BUbiNG LAW-Unimi Java

The LAW next generation crawler.

92
MediaCrawler
MediaCrawler RaidenEI21 Python

MediaCrawler is a powerful web scraper for self-media platforms. Easily collect and analyze content to enhance your digital strategy. 🌐🕷️

92
MedicalKG
MedicalKG yeeeqichen Python

医疗知识图谱构建实战,通过爬虫获取百度百科数据,使用Mongodb存储结构化三元组,并使用neo4j进行知识图谱的构建及可视化; Medical Knowledge Graph; Crawler;...

92
HydraRecon
HydraRecon aufzayed Python

All In One, Fast, Easy Recon Tool

92
Amazon-Price-Alert
Amazon-Price-Alert GaryniL Python

Price tracker of Amazon

91
SeleniumDemo
SeleniumDemo tobecrazy HTML

Selenium automation test framework

90
crawlie
crawlie nietaki Elixir

A simple Elixir library for writing decently-performing crawlers with minimum effort.

90
Novel-crawler
Novel-crawler ling7334 Python

这是一个用Python写的小说爬虫软件

90
html-table-extractor
html-table-extractor yuanxu-li Python

extract data from html table

89
movie-elasticsearch
movie-elasticsearch cbwleft Java

使用 SpringBoot2.0+ElasticSearch 实现的开源电影搜索引擎

89
scrapeGPT
scrapeGPT LexiestLeszek Python

ScrapeGPT is a RAG-based Telegram bot designed to scrape and analyze websites, then answer questions based on the scraped content. The bot utilizes Re...

88
shopify-spy
shopify-spy ndgigliotti Python

Extract structured data from Shopify websites.

88
WebScraper
WebScraper MLArtist Python

Python-based web crawling script with randomized intervals, user-agent rotation, and proxy server IP rotation to outsmart website bots and prevent blo...

88
WebSecurityArticles
WebSecurityArticles zongdeiqianxing Python

爬取及整理Freebuf\安全客\先知\知道创宇等站点的”web安全“类优质文章

87