Topic

crawler

Repositories (1470)

google-play-scraper
google-play-scraper JoMingyu Python

Google play scraper for Python inspired by <facundoolano/google-play-scraper>

1k
stormcrawler
stormcrawler apache Java

A scalable, mature and versatile web crawler based on Apache Storm

995
SpiderSuite
SpiderSuite spidersuite

SpiderSuite (web security crawler) releases, wiki and roadmap

979
magnet-dht
magnet-dht chenjiandongx Python

✌️ Python3 BitTorrent DHT crawler

965
crawler
crawler fredwu Elixir

A high performance web crawler / scraper in Elixir.

956
x-kit
x-kit xiaoxiunique Python

Pure-Python protocol tools and research notes for X.com, with VibeLoft Twitter account datasets

940
SecCrawler
SecCrawler Le0nsec Go

一个方便安全研究人员获取每日安全日报的爬虫和推送程序,目前爬取范围包括先知社区、安全客、Seebug Paper、跳跳糖、奇安信攻防社区、棱角社区以及绿盟、腾讯玄...

939
icrawler
icrawler hellock Python

A multi-thread crawler framework with many builtin image crawlers provided.

932
spider_reverse
spider_reverse 0xAllenChen Python

爬虫逆向案例,已完成:TLS指纹|瑞数|震坤行 | 网易易盾 | 微信小程序反编译逆向(百达星系) | 同花顺 | rpc解密 | 加速乐 | 极验滑块验证码 | 巨量算数 | Boss...

926
zhihu-crawler
zhihu-crawler wycm Java

zhihu-crawler是一个基于Java的高性能、支持免费http代理池、支持横向扩展、分布式爬虫项目

921
chatWeb
chatWeb SkywalkerDarren Python

ChatWeb can crawl web pages, read PDF, DOCX, TXT, and extract the main content, then answer your questions based on the content, or summarize the key...

916
BaiduImageSpider
BaiduImageSpider kong36088 Python

一个超级轻量的百度图片爬虫

913
TumblThree
TumblThree johanneszab C#

A Tumblr Blog Backup Application

908
siteone-crawler
siteone-crawler janreges Rust

SiteOne Crawler is a cross-platform website crawler and analyzer for SEO, security, accessibility, and performance optimization—ideal for developers,...

906
TikHub-API-Python-SDK
TikHub-API-Python-SDK TikHub Python

High-performance asynchronous Douyin(抖音) TikTok Xiaohongshu(小红书) Kuaishou(快手) Weibo(微博) Instagram YouTube(油管) Twitter(X) Captcha Solver(验...

887
scrapyrt
scrapyrt scrapinghub Python

HTTP API for Scrapy spiders

884
skrape.it
skrape.it skrapeit Kotlin

A Kotlin-based testing/scraping/parsing library providing the ability to analyze and extract data from HTML (server & client-side rendered). It places...

875
bookcorpus
bookcorpus soskek Python

Crawl BookCorpus

864
ArrowDL
ArrowDL setvisible C++

ArrowDL (Arrow Downloader) is a download manager for Windows, MacOS and Linux

848
Weibo-Analyst
Weibo-Analyst KimMeen Python

Social media (Weibo) comments analyzing toolbox in Chinese 微博评论分析工具, 实现功能: 1.微博评论数据爬取; 2.分词与关键词提取; 3.词云与词频统计; 4.情...

838
spidr
spidr postmodern Ruby

A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to...

836
Scavenger
Scavenger rndinfosecguy Python

Crawler (Bot) searching for credential leaks on paste sites.

829
easy-scraping-tutorial
easy-scraping-tutorial MorvanZhou Jupyter Notebook

Simple but useful Python web scraping tutorial code.

823
awesome-ai-reverse
awesome-ai-reverse darbra

ai reverse 一把梭

817
till
till DataHenHQ Go

DataHen Till is a companion tool to your existing web scraper that instantly makes it scalable, maintainable, and more unblockable, with minimal code...

815
seo-audits-toolkit
seo-audits-toolkit StanGirard Python

SEO & Security Audit for Websites. Lighthouse & Security Headers crawler, Sitemap/Keywords/Images Extractor, Summarizer, etc ...

812
course-crawler
course-crawler Foair Python

🎓 中国大学MOOC、学堂在线、网易云课堂、好大学在线、爱课程 MOOC 课程下载。

810
jvppeteer
jvppeteer fanyong920 Java

Java API For Chrome and Firefox

805
xeHentai
xeHentai fffonion Python

Doujinshi downloader 绅士漫画下载

803
pic-gather
pic-gather Licoy

🛑 image collector, which supports custom acquisition source configuration and is compatible with MacOS and Windows operating systems.

801
Lulu
Lulu iawia002 Python

[Unmaintained] A simple and clean video/music/image downloader 👾

800
fetchbot
fetchbot PuerkitoBio Go

A simple and flexible web crawler that follows the robots.txt policies and crawl delays.

790
js-cookie-monitor-debugger-hook
js-cookie-monitor-debugger-hook JSREI TypeScript

js cookie逆向利器:js cookie变动监控可视化工具 & js cookie hook打条件断点

784
seonaut
seonaut StJudeWasHere Go

Open source SEO audit tool.

783
creeper
creeper wspl Go

:paw_prints: Creeper - The Next Generation Crawler Framework (Go)

777
linkedin-profile-scraper-api
linkedin-profile-scraper-api josephlimtech TypeScript

🕵️‍♂️ LinkedIn profile scraper returning structured profile data in JSON.

776
lxBook
lxBook lixi5338619 JavaScript

《爬虫逆向进阶实战》书籍代码库

769
BaiduSpider
BaiduSpider BaiduSpider Python

BaiduSpider,一个爬取百度搜索结果的爬虫,目前支持百度网页搜索,百度图片搜索,百度知道搜索,百度视频搜索,百度资讯搜索,百度文库搜索,百度经验搜索和百...

765
TheBigBrother
TheBigBrother chadi0x Python

The Big Brother V6.0 is a weaponized OSINT platform featuring username enumeration (473+ platforms), quad-vector visual intelligence, Sky Radar tracki...

761
xxl-crawler
xxl-crawler xuxueli Java

A lightweight web crawler framework.(Java爬虫框架)

758
hacker-news-digest
hacker-news-digest polyrabbit Python

:newspaper: Let ChatGPT Summarize Hacker News for You

756
Krawl
Krawl BlessedRebuS Python

Krawl is a customizable, lightweight, cloud-native web deception server and anti-crawler that creates fake web applications with low-hanging vulnerabi...

738
TumblThree
TumblThree TumblThreeApp C#

A Tumblr and Twitter Blog Backup Application

734
PyPtt
PyPtt PyPtt Python

The best PTT library

732
wscan
wscan chushuai Go

Wscan is a web security scanner that focuses on web security, dedicated to making web security accessible to everyone.

713
Free_Proxy_Website
Free_Proxy_Website cyubuchen Python

获取免费socks/https/http代理的网站集合

700
fbcrawl
fbcrawl rugantio Python

A Facebook crawler

694
gOSINT
gOSINT Nhoya Go

OSINT Swiss Army Knife

675
Search-Engines-Scraper
Search-Engines-Scraper tasos-py Python

Search google, bing, yahoo, and other search engines with python

671
FileMasta
FileMasta ohhsodead C#

A search application to explore, discover and share online files

670