Topic

crawler

Repositories (1470)

crawler
crawler infinilabs Go

🕷️ An easy-to-use spider written in Golang. (previous named GOPA.)

311
gplay-scraper
gplay-scraper Mohammedcha Python

GPlay Scraper is a powerful Python Google Play scraper library for extracting comprehensive app data from the Google Play Store. Scrape Google Play St...

309
Python-Web-Scraping-Tutorial
Python-Web-Scraping-Tutorial oxylabs Python

In this Python Web Scraping Tutorial, we will outline everything needed to get started with web scraping. We will begin with simple examples and move...

307
SpotifyScraper
SpotifyScraper AliAkhtari78 Python

Extract public Spotify data — tracks, albums, artists, playlists, podcasts & lyrics — without the official API. Sync + async, typed models, one depend...

305
Fast-LianJia-Crawler
Fast-LianJia-Crawler CaoZ Python

直接通过链家 API 抓取数据的极速爬虫,宇宙最快~~ 🚀

303
js-reverse
js-reverse freedom-wy HTML

JS逆向研究

303
xvideos
xvideos rodrigogs TypeScript

xvideos API library

302
site-audit-seo
site-audit-seo viasite JavaScript

Web service and CLI tool for SEO site audit: crawl site, lighthouse all pages, view public reports in browser. Also output to console, json, csv, xlsx

302
bitextor
bitextor bitextor Python

Bitextor generates translation memories from multilingual websites

300
weibo-topic-spider
weibo-topic-spider czy1999 Python

微博超级话题爬虫,微博词频统计+情感分析+简单分类,新增肺炎超话爬取数据

299
line-bot-tutorial
line-bot-tutorial twtrubiks Python

line-bot-tutorial use python flask

296
Sasila
Sasila da2vin Python

一个灵活、友好的爬虫框架

294
ComicCrawler
ComicCrawler eight04 Python

An image crawler written in Python.

294
pychromeless
pychromeless jairovadillo Python

Python Lambda Chrome Automation (naming pending)

292
PixivCrawler
PixivCrawler cwher Python

Pixiv Utils implemented in Python, including Pixiv Crawler and Mosaic Puzzles, support for rankings, personal bookmarks, artist works and keyword sear...

291
Gorecon
Gorecon devanshbatham Go

Gorecon is a All in one Reconnaissance Tool , a.k.a swiss knife for Reconnaissance , A tool that every pentester/bughunter might wanna consider into...

286
LinkedIn-Scraper
LinkedIn-Scraper TufayelLUS Python

A LinkedIn Scraper to scrape up to 1k LinkedIn profiles(due to LinkedIn limit) from company profile links and save their e-mail addresses if available...

285
Instagram-Bot
Instagram-Bot mustafadalga Python

An Instagram bot developed using the Selenium Framework

284
th-music-video-generator
th-music-video-generator Jasonnor JavaScript

Touhou Project random music video generator/player, crawling image and video from websites to generate MV.

280
Strong-Web-Crawler
Strong-Web-Crawler microfisher C#

基于C#.NET+PhantomJS+Sellenium的高级网络爬虫程序。可执行Javascript代码、触发各类事件、操纵页面Dom结构。

276
go-movies
go-movies hezhizheng Go

golang spider Crawler 爬虫 电影

275
laravel-seo-scanner
laravel-seo-scanner backstagephp PHP

Scan your Laravel application routes for SEO improvements suggestions.

272
weiboPicDownloader
weiboPicDownloader yAnXImIN Java

免登录下载微博图片 爬虫 Download Weibo Images without Logging-in

271
D4N155
D4N155 OWASP Shell

OWASP D4N155 - Intelligent and dynamic wordlist using OSINT

271
onecomic
onecomic hardwarecode Python

一本漫画

271
rotating-tor-http-proxy
rotating-tor-http-proxy zhaow-de Shell

A multi-arch image provides one HTTP proxy endpoint with many concurrent tunnels to the Tor network.

267
crypto-crawler-rs
crypto-crawler-rs crypto-crawler Rust

A rock-solid cryptocurrency crawler library.

267
algoliasearch-netlify
algoliasearch-netlify algolia TypeScript

Official Algolia Plugin for Netlify. Index your website to Algolia when deploying your project to Netlify with the Algolia Crawler

265
antch
antch antchfx Go

Antch, a fast, powerful and extensible web crawling & scraping framework for Go

265
RuiJi.Net
RuiJi.Net zhupingqi C#

crawler framework, distributed crawler extractor

264
ptt-alertor
ptt-alertor Ptt-Alertor Go

:loudspeaker: Ptt 文章通知機器人!Notify Ptt Article in Realtime

263
galer
galer dwisiswant0 Go

A fast tool to fetch URLs from HTML attributes by crawl-in.

263
Github-spider
Github-spider chenjiandongx Python

Github 仓库及用户分析爬虫

260
weibo_terminator_workflow
weibo_terminator_workflow lucasjinreal Python

Update Version of weibo_terminator, This is Workflow Version aim at Get Job Done!

258
Selenops
Selenops zntfdr Swift

A Swift Web Crawler 🕷

258
robots-txt
robots-txt spatie PHP

Determine if a page may be crawled from robots.txt, robots meta tags and robot headers

258
arachnid
arachnid zrashwani PHP

Crawl all unique internal links found on a given website, and extract SEO related information - supports javascript based sites

254
FileSensor
FileSensor Xyntax Python

Dynamic file detection tool based on crawler 基于爬虫的动态敏感文件探测工具

253
ok_ip_proxy_pool
ok_ip_proxy_pool cwjokaka Python

🍿爬虫代理IP池(proxy pool) python🍟一个还ok的IP代理池

253
FindJobs-Agent
FindJobs-Agent he-yufeng Python

LLM-powered toolkit for skill analysis, AI interviews, resume scoring, and job structuring. Automates professional skill taxonomy and interview proces...

252
InfinityCrawler
InfinityCrawler TurnerSoftware C#

A simple but powerful web crawler library for .NET

252
CSharpCrawler
CSharpCrawler zhaotianff C#

C#爬虫示例程序,想学习爬虫入门知识的可以看过来。后续会慢慢加入更多爬虫相关的知识。

251
chromium_for_spider
chromium_for_spider myvyang HTML

dynamic crawler for web vulnerability scanner

250
siteone-crawler-gui
siteone-crawler-gui janreges Svelte

SiteOne Crawler GUI is a cross-platform website crawler and analyzer for SEO, security, accessibility, and performance optimization—ideal for develope...

250
Sub
Sub Leon406 Kotlin

节点爬取,筛选, 支持Clash,base64订阅解析,自动生成可用的ss, ssr, v2ray, trojan节点. 已集成Github Action,每天8-24,定时更新.

249
ZhihuSpider
ZhihuSpider kong36088 Python

多线程知乎用户爬虫,基于python3

247
gf-secrets
gf-secrets dwisiswant0 Shell

Secret and/or credential patterns used for gf.

246
woid
woid vitorfs Python

Simple news aggregator displaying top stories in real time

245
Sitemap-Generator-Crawler
Sitemap-Generator-Crawler vezaynk PHP

PHP script to recursively crawl websites and generate a sitemap. Zero dependencies.

245
Tumblr_Crawler
Tumblr_Crawler sparrow629 Python

This is a Multi-thread crawler for Tumblr.

244