Topic

crawler

Repositories (1456)

damai-tickets
damai-tickets Jxpro Python

大麦抢票脚本案例

128
lumberjack
lumberjack JakePartusch JavaScript

An automated website accessibility scanner and cli

126
dyer
dyer hominee Rust

Dyer is designed for reliable, flexible and fast web crawling, providing some high-level, comprehensive features without compromising speed.

126
instagram-profilecrawl
instagram-profilecrawl nacimgoura JavaScript

:computer: Quickly crawl the information (e.g. followers, tags, etc...) of an instagram profile. No login required!

126
GoogleImagesDownloader
GoogleImagesDownloader WuLC Python

Enlarge training dataset by searching images with specified keywords in google and download the presented images

126
Skill-Share-Crawler---DL
Skill-Share-Crawler---DL tharyckgusmao JavaScript

Download Videos Skill Share per ID or per Class

125
sentinel-crawler
sentinel-crawler wx-chevalier JavaScript

Xenomorph Crawler, a Concise, Declarative and Observable Distributed Crawler(Node / Go / Java / Rust) For Web, RDB, OS, also can act as a Monitor(with...

125
prerender-java
prerender-java greengerong Java

java framework for prerender

125
diffbot-python
diffbot-python diffbot Python

Python client library for Diffbot APIs

124
graphquery
graphquery storyicon Go

GraphQuery is a query language and execution engine tied to any backend service.

124
pku3b
pku3b sshwy Rust

🎓a Better BlackBoard for PKUers. 北京大学教学网命令行工具(🖥️Win/🐧Linux/🍏Mac)

123
Price-monitor
Price-monitor qqxx6661 Python

某东商品价格监控:自定义商品价格,降价邮件/微信提醒。技术:Python爬虫/IP代理池/JS接口爬取/Selenium页面爬取

122
aiotieba
aiotieba Starry-OvO Python

百度贴吧吧务管理器✨删帖机✨使用aiohttp封装大量贴吧核心API

122
TiebaManager
TiebaManager xfgryujk C++

(已跑路)百度贴吧吧务管理工具,自动扫描帖子并处理违规帖

122
andvaranaut
andvaranaut glouw C

A dungeon crawler

122
goClone
goClone shurco Go

🌱 goClone - clone websites in seconds

121
AmazonRobot
AmazonRobot WuLC Python

Amazon商品引流的 python 爬虫

121
findpapers
findpapers jonatasgrosman Python

Findpapers: A tool for helping researchers who are looking for related works

120
eyes
eyes r05323028 Python

Public Opinion Mining System of Taiwanese Forums

119
Lcrawl
Lcrawl lndj PHP

一只优雅的正方教务系统爬虫。

119
BaiduCrawler
BaiduCrawler mazzzystar Python

Sample of using proxies to crawl baidu search results.

118
bots-zoo
bots-zoo antoinevastel JavaScript
117
ThesaurusSpider
ThesaurusSpider WuLC Python

下载搜狗、百度、QQ输入法的词库文件的 python 爬虫,可用于构建不同行业的词汇库

117
proxy-pool
proxy-pool denghuichao Java

爬虫代理IP池服务,可供其他爬虫程序通过restapi获取

116
LFITester
LFITester kostas-pa Python

LFITester is a Python3 program that automates the detection and exploitation of Local File Inclusion (LFI) vulnerabilities on a server.

115
starfish-ql
starfish-ql SeaQL Rust

✴️ An experimental graph database

114
sitemap-extract
sitemap-extract phase3dev Python

Processes XML sitemaps and extracts URLs. Includes features such as support for both plain XML and compressed XML files, multiple input sources, prote...

113
linkcrawler
linkcrawler schollz Go

Cross-platform persistent and distributed web crawler :link:

113
scrapai-cli
scrapai-cli discourselab Python

AI-powered web scraping CLI. Describe what you want, get a production-ready Scrapy spider. Write once, reuse forever.

113
APSoft-Web-Scanner-v2
APSoft-Web-Scanner-v2 ph09nix C#

Powerful dork searcher and vulnerability scanner for windows platform

112
wxpath
wxpath rodricios Python

wxpath - declarative web crawling with XPath; a Web Query Language (WQL)

112
Crawler
Crawler phantom-sea-limited Python

针对某亿些小说网站的爬虫

112
pagser
pagser foolin Go

Pagser is a simple, extensible, configurable parse and deserialize html page to struct based on goquery and struct tags for golang crawler

111
ManyACG
ManyACG krau Go

Collect, Download, Organize and Share your Favorite Anime Artworks.

111
goscraper
goscraper badoux Go

Golang pkg to quickly return a preview of a webpage (title/description/images)

111
xcrawl3r
xcrawl3r hueristiq Go

A command-line utility designed to recursively spider webpages for URLs. It works by actively traversing websites - following links embedded in webpag...

111
bee-university
bee-university beecost Python

Project thu thập điểm chuẩn đại học 2014 - 2018 và phân tích dữ liệu

111
Proxy-List-Scrapper
Proxy-List-Scrapper narkhedesam Python

Proxy List Scrapper

110
qcrawl
qcrawl crawlcore Python

qcrawl - fast async web crawling & scraping framework for Python.

110
scrapy-puppeteer
scrapy-puppeteer clemfromspace Python

Scrapy + Puppeteer

110
gocrawler
gocrawler superjcd Go

gocrawler, go分布式爬虫框架

109
spider-py
spider-py spider-rs Rust

Spider ported to Python

109
zyte-smartproxy-headless-proxy
zyte-smartproxy-headless-proxy zytedata Go

A complimentary proxy to help to use SPM with headless browsers

109
antispider
antispider dytttf JavaScript
109
crawler
crawler brantou Python

爬虫, http代理, 模拟登陆!

107
images-web-crawler
images-web-crawler amineHorseman Python

This package is a complete tool for creating a large dataset of images (specially designed -but not only- for machine learning enthusiasts). It can cr...

107
weibo-scraper
weibo-scraper Xarrow Python

Simple Weibo Scraper

107
bose
bose omkarcloud Python

✨ BOSE IS SWISS ARMY KNIFE 🔪 FOR BOT DEVELOPMENT. THE ULTIMATE BOT DEVELOPMENT FRAMEWORK. 🤖

107
aliexscrape
aliexscrape ducdev JavaScript

Get Aliexpress product details in JSON

106
pappet
pappet patrickschur JavaScript

A command-line tool to crawl websites using puppeteer.

105