Topic

crawler

Repositories (1456)

indonesian-NLP-resources
indonesian-NLP-resources kirralabs

data resource untuk NLP bahasa indonesia

231
crawlab-lite
crawlab-lite crawlab-team Vue

Lite version of Crawlab. 轻量版 Crawlab 爬虫管理平台

230
goose-parser
goose-parser redco JavaScript

Universal scraping tool, which allows you to extract data using multiple environments

229
facebook-data-extraction
facebook-data-extraction 18520339 Python

Experience for effectively fetching Facebook data by Querying Graph API with Account-based Token and Operating undetectable scraping Bots to extract C...

229
darc
darc JarryShaw Python

Darkweb Crawler Project

227
KoreaNewsCrawler
KoreaNewsCrawler lumyjuwon Python

A korean news crawler built to ingest large amounts of news data.

225
black-widow
black-widow offensive-hub Python

GUI based offensive penetration testing tool (Open Source)

225
google-group-crawler
google-group-crawler icy Shell

[Deprecated] Get (almost) original messages from google group archives. Your data is yours.

224
WebVideoBot
WebVideoBot tim232385 Java

Web crawler.

223
FooProxy
FooProxy 01ly Python

稳健高效的评分制-针对性- IP代理池 + API服务,可以自己插入采集器进行代理IP的爬取,针对你的爬虫的一个或多个目标网站分别生成有效的IP代理数据库,支持Mongo...

222
N2H4
N2H4 forkonlp R

네이버 뉴스 수집을 위한 도구

222
weibo_wordcloud
weibo_wordcloud gaussic Python

根据关键词抓取微博数据,再生成词云

221
scrapy-zhihu-github
scrapy-zhihu-github zhijunio Python

scrapy examples for crawling zhihu and github

221
MetaFinder
MetaFinder Josue87 Python

Search for documents in a domain through Search Engines (Google, Bing and Baidu). The objective is to extract metadata

221
91porn-crawler
91porn-crawler blue-troy Java

91 porn crawler. 自动爬取并下载你想要的91porn热门视频。Automatically download your "favorite" 91porn hot movies.

220
proxyhub
proxyhub ForceFledgling Python

An advanced [Finder | Checker | Server] tool for proxy servers, supporting both HTTP(S) and SOCKS protocols. 🎭

216
scrapedin-linkedin-crawler
scrapedin-linkedin-crawler linkedtales JavaScript

Crawler for LinkedIn full profiles 2019

215
JavPy
JavPy TheodoreKrypton JavaScript

Enjoy driving on a Javascriptive (originally Pythonic) way to Japanese AV!

214
goribot
goribot zhshch2002 Go

[Crawler/Scraper for Golang]🕷A lightweight distributed friendly Golang crawler framework.一个轻量的分布式友好的 Golang 爬虫框架。

211
WebScrapper
WebScrapper nuhmanpk Python

Powerful Telegram bot for web scraping and crawling. Fast, easy, and loved by thousands!

208
scrapemate
scrapemate gosom Go

Golang Crawling and scraping framework

205
crawler
crawler Norconex Java

Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to v...

204
ChainWalker
ChainWalker 0xsha Go

Rapid Smart Contract Crawler

204
awesome-python-primer
awesome-python-primer zkqiang Python

自学入门 Python 优质中文资源索引,包含 书籍 / 文档 / 视频,适用于 爬虫 / Web / 数据分析 / 机器学习 方向

203
laosj
laosj songtianyi Go

golang light-weight image crawler

202
instagram-crawler
instagram-crawler mgleon08 Ruby

Crawl instagram photos, posts and videos for download.

202
NoSmoke
NoSmoke macacajs JavaScript

A cross platform UI crawler which scans view trees then generate and execute UI test cases.

202
authority-data
authority-data yiliyassh Python

官方权威数据:统计年签,统计公报,互联网行业报告,工信部数据,ICT报告等 Official authoritative data (Chinese)

201
packagist-mirror
packagist-mirror webysther PHP

📦✂️📋📦 Create a mirror of packagist.org metadata for use locally with composer

200
VideoServer
VideoServer GF-Allen JavaScript

以Node.js基于express以及爬虫实现的视频资源后端

199
slrp
slrp nfx Go

rotating open proxy multiplexer

199
search
search subins2000 PHP

An Open Source Search Engine

198
pkulaw_spider
pkulaw_spider yinhao0214 Python

爬取北大法宝网http://www.pkulaw.cn/Case/

198
ghs
ghs seart-group Java

GitHub Search: Platform used to crawl, store and present projects from GitHub, as well as any statistics related to them

197
NewsCrawler
NewsCrawler Jacen789 Python

新闻爬虫,爬取新浪、搜狐、新华网即时财经新闻。

196
gflare-tk
gflare-tk beb7 Python

Open-Source Python Based SEO Web Crawler

195
cocrawler
cocrawler cocrawler Python

CoCrawler is a versatile web crawler built using modern tools and concurrency.

194
web-bee
web-bee codesofun Java

🐝 Web vertical crawler framework for fun

193
zhihu-crawler-people
zhihu-crawler-people elliotxx Python

A simple distributed crawler for zhihu && data analysis

193
ir-search
ir-search djfksjd Python

🇰🇷 한국 정부 지원사업 전수조사 에이전트 스킬 [Claude Code·Codex·agy(Antigravity)·Cursor·Gemini CLI·Grok Build 지원] K-Startup·기업마당·NIPA·KOCCA·SMTE...

193
digger
digger hetianyi Go

Digger is a powerful and flexible web crawler implemented by pure golang

191
zhihu_fun
zhihu_fun AnyISalIn JavaScript

基于 Selenium 的知乎关键词爬虫

186
leetcode-spider
leetcode-spider Ma63d JavaScript

用 node.js 爬你自己的 leetcode 解题源码

186
crawler-for-github-trending
crawler-for-github-trending poozhu JavaScript

🕷️ A node crawler for github trending.

185
kuaishou-crawler
kuaishou-crawler oGsLP Python

As you can see, a kuaishou crawler

185
nCov2019_data_crawler
nCov2019_data_crawler LiuTianyong Python

疫情数据爬虫,2019新型冠状病毒数据仓库,轨迹数据,同乘数据,报道

184
gogetcrawl
gogetcrawl karust Go

Extract web archive data using Wayback Machine and Common Crawl

184
telegram-groups-crawler
telegram-groups-crawler edogab33 Python

A Telegram crawler made in Python to automatically search groups and channels and collect any type of data from them.

184
sensitivefilescan
sensitivefilescan aipengjie Python
183
datmusic-api
datmusic-api alashow PHP
182