Topic

crawler

Repositories (1456)

WeiBoCrawler
WeiBoCrawler zhouyi207 Rust

微博数据采集,后续会加上知乎,贴吧,小红书,抖音,快手等主流媒体内容

149
not-your-average-web-crawler
not-your-average-web-crawler tijme Python

A web crawler (for bug hunting) that gathers more than you can imagine.

149
pylinkvalidator
pylinkvalidator bartdag Python

pylinkvalidator is a standalone and pure python link validator and crawler that traverses a web site and reports errors (e.g., 500 and 404 errors) enc...

149
GoodreadsScraper
GoodreadsScraper havanagrawal Python

Scrape data from Goodreads using Scrapy and Selenium :books:

148
pinscrape
pinscrape iamatulsingh Python

A simple library to scrape Pinterest images.

147
jlitespider
jlitespider luohaha Java

A lite distributed Java spider framework :-)

147
Web-Data-Scraper
Web-Data-Scraper umbrellaDocumentation JavaScript

Web Data Scraper - no-code internet scraping. Extract and export to CSV, Excel, JSON, Google Sheets, and Webhook.

147
crawler_detect
crawler_detect loadkpi Ruby

Ruby gem to detect bots and crawlers via the user agent

147
pachong
pachong jin10086 Jupyter Notebook

一些爬虫的代码

146
moe-copy-ai
moe-copy-ai yusixian TypeScript

✨ 萌萌哒的 AI 网页数据提取助手 ✨

145
TiebaArchiver
TiebaArchiver Sorceresssis Python

保存百度贴吧帖子到本地,并且支持图片, 视频, 语音等内容。与本项目配套的阅读器 TiebaReader(https://github.com/Sorceresssis/TiebaReader)

145
sasori
sasori karthikuj JavaScript

Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.

145
bilibili_member_crawler
bilibili_member_crawler cwjokaka Python

B站用户爬虫 好耶~是爬虫

143
npm-search
npm-search algolia TypeScript

🗿 npm ↔️ Algolia replication tool :skier: :snail: :artificial_satellite:

143
agent-line-bot
agent-line-bot Lin-jun-xiang Python

🤖Free Agent Line Bot with Web Search, Google Image Search, Image Generator, Video Generator...

142
WebReaper
WebReaper alex-on-ai C#

AI-native web scraper. Single binary with a bundled Claude Code skill. MIT-licensed alternative to Firecrawl.

142
pixiv_func_mobile
pixiv_func_mobile git-xiaocao Dart

功能齐全的Pixiv第三方客户端 免代理 支持查看动图查看小说

142
auto-lighthouse
auto-lighthouse TGiles HTML

A utility package for automating lighthouse reporting

142
convertible-bond-crawler
convertible-bond-crawler jackluson HTML

宁稳网(旧富投网)、集思录可转债数据&策略分析

140
PHPCreeper
PHPCreeper blogdaren PHP

A new generation of multi-process async event-driven spider engine based on workerman. Support headless browser. 🌿基于workerman实现的多进程异步事件...

140
poopak
poopak teal33t Python

POOPAK - TOR Hidden Service Crawler

140
KamiYomu
KamiYomu KamiYomu C#

A self-hosted, extensible manga reader and download tool with plug-in support.

139
Ceiba-Downloader
Ceiba-Downloader jameshwc Python

This is a course-downloader to help NTU students download courses data from NTU Ceiba.

139
docs
docs zhangslob

《数据采集从入门到放弃》源码。内容简介:爬虫介绍、就业情况、爬虫工程师面试题 ;HTTP协议介绍; Requests使用 ;解析器Xpath介绍; MongoDB与MySQL; 多线程...

139
taki
taki egoist TypeScript

Take a snapshot of any website.

139
pricetrack
pricetrack duyet JavaScript

Price tracker monitors of products and alerts you when prices drop. Supported tiki.vn, shopee, lotte.vn, ... Built with firebase https://pricetrack.we...

139
WeiboSpider
WeiboSpider CharesFang Python

微博爬虫,一个基于Scrapy框架的轻量微博爬虫,Sina Weibo Spider

138
sitemapper
sitemapper seantomburke TypeScript

Parse through any sitemap in Node.js

137
wget-lua
wget-lua ArchiveTeam C

Wget-AT is a modern Wget with Lua hooks, Zstandard (+dictionary) WARC compression and URL-agnostic deduplication.

137
fuckBookWalker
fuckBookWalker VermiIIi0n Python

Download books from bookwalker.jp/bookwalker.com.tw

137
PatentCrawler
PatentCrawler will4906 Python

scrapy专利爬虫(停止维护)

137
GoodBots
GoodBots AnTheMaker

Updated lists of IP addresses/whitelists of good bots and crawlers. Includes GoogleBot, BingBot, DuckDuckBot, etc.

136
php-crawler
php-crawler hedii PHP

A php crawler that finds emails on the internets

136
blinkist-m4a-downloader
blinkist-m4a-downloader luckylittle Go

Grabs all of the audio files from all of the Blinkist books

136
spiderbuf
spiderbuf hhuayuan Python

Spiderbuf 是一个专注于 Python 爬虫练习的网站。提供丰富的爬虫教程、爬虫案例解析和爬虫练习题。Python爬虫开发强化练习,在矛与盾的攻防中不断提高技术水平,...

135
proxy-pool
proxy-pool XiaomingX Python

Python ProxyPool for web spider.ProxyPool 是一个用于采集、验证和管理代理IP的轻量级工具,旨在帮助用户自动维护高质量的代理池,方便在爬虫、网络请求中灵活...

134
Terpene-Profile-Parser-for-Cannabis-Strains
Terpene-Profile-Parser-for-Cannabis-Strains MaxValue Python

Parser and database to index the terpene profile of different strains of Cannabis from online databases

134
leetcode-ranking-search
leetcode-ranking-search chiehmin Vue

Leetcode Contest Ranking Searcher

134
picacomic_downloader
picacomic_downloader muyoou Python

哔咔漫画收藏夹下载程序

134
ParseHub
ParseHub z-mio Python

轻量、异步、开箱即用的社交媒体聚合解析库

133
feaplat
feaplat Boris-code

爬虫管理系统,支持集群,弹性伸缩。支持运行feapder、scrapy、selenium、playwright等各种框架及脚本

132
news-crawler
news-crawler LuChang-CS Python

A news crawler for BBC News, Reuters and New York Times.

132
web-scout-mcp
web-scout-mcp pinkpixel-dev JavaScript

A powerful MCP server extension providing web search and content extraction capabilities. Integrates DuckDuckGo search functionality and URL content e...

130
Sina-Weibo-Album-Downloader
Sina-Weibo-Album-Downloader lincanbin Python

Multithreading download all HD photos / pictures from someone's Sina Weibo album.

130
node-crawler
node-crawler ethereum Go

Attempts to crawl the Ethereum network of valid Ethereum execution nodes and visualizes them in a nice web dashboard.

130
learncpp-download
learncpp-download amalrajan Python

Multi-threaded web scraper to download all the tutorials from www.learncpp.com and convert them to PDF files concurrently.

129
scraply
scraply alash3al Go

Scraply a simple dom scraper to fetch information from any html based website

129
pdf-crawler
pdf-crawler SimFin Python

SimFin's open source PDF crawler

129
onegram
onegram pauloromeira Python

This repository is no longer maintained.

128
damai-tickets
damai-tickets Jxpro Python

大麦抢票脚本案例

128