Topic

crawler

Repositories (1460)

talospider
talospider howie6879 Python

talospider - A simple,lightweight scraping micro-framework

57
All-IT-eBooks-Spider
All-IT-eBooks-Spider Kulbear Python

[Updated] A simple python crawler for my tutorial blog at http://www.jianshu.com/p/8fb5bc33c78e

57
instagram-hashtag-crawler
instagram-hashtag-crawler simonseo Python

Crawl Instagram hashtags

57
wishlist
wishlist Jaymon Python

Read an Amazon wishlist programmatically with Python

57
crawler
crawler tomasnorre PHP

Libraries and scripts for crawling the TYPO3 page tree. Used for re-caching, re-indexing, publishing applications etc.

57
PicCrawler
PicCrawler fengzhizi715 Java

使用RxJava2 和 Java 8的特性开发的图片爬虫

56
devsearch
devsearch nicholaskajoh Python

A web search engine built with Python which uses TF-IDF and PageRank to sort search results.

56
price-monitoring
price-monitoring roccomuso JavaScript

Node.js price monitoring library, leveraging the power of x-ray and nightmare.

56
actor-facebook-scraper
actor-facebook-scraper pocesar TypeScript

Scrape public Facebook pages, posts, reviews and comments

56
crawler
crawler a11ywatch Rust

gRPC web crawler turbo charged for performance

56
mtywatch
mtywatch jufeng-2022

一句话监控网页内容变化,AI | 爬虫 | 网页监控 | 网页更新提醒 | 网页内容订阅

56
wechat_biz
wechat_biz yusp998 Go

微信公众号爬虫,以API方式提供公众号文章获取,包括阅读量、点赞等

56
NewsCrap
NewsCrap odaysec Python

NewsCrap adalah alat scraping berita Google berbasis Command Line Interface (CLI) yang dirancang untuk riset, investigasi, dan pengumpulan data OSINT....

56
MyCrawler
MyCrawler netcan Python

我的爬虫合集

55
Deepminer
Deepminer Conso1eCowb0y Python

Deep web crawler and search engine

55
rag-crawler
rag-crawler sigoden TypeScript

Crawl a website to generate knowledge file for RAG

55
kalel
kalel noobscode Python

Kal El Network Stress Test and Penetration Testing Toolkit

54
browser-as-a-service
browser-as-a-service hfreire JavaScript

A web browser :earth_americas: hosted as a service, to render your JavaScript web pages as HTML

54
crawler_shopee
crawler_shopee charlie0227 Python

Shopee coin getter is a script to collect daily shopee coins.

54
chan-downloader
chan-downloader mariot Rust

CLI to download all images/webms in a 4chan thread

54
rarbgcli
rarbgcli FarisHijazi Python

RARBG command line interface for scraping the rarbg.to torrent search engine

54
uforall
uforall rix4uni Go

uforall is a fast url crawler this tool crawl all URLs number of different sources, alienvault,WayBackMachine,urlscan,commoncrawl

54
crawlkit
crawlkit openclaw Go

Shared Go infrastructure for local-first crawler archives.

54
fb-page-chat-download
fb-page-chat-download eisenjulian Python

Python script to download messages from a Facebook page to a CSV file

53
flink-crawler
flink-crawler kkrugler Java

Continuous scalable web crawler built on top of Flink and crawler-commons

53
PageParser
PageParser mouday Python

网页解析器,用于网络爬虫解析页面, 不懂网页解析也能写爬虫

53
python-scrapfly
python-scrapfly scrapfly Python

Scrapfly Python SDK for headless browsers and proxy rotation

53
subscan
subscan eredotpkfr Rust

⚡ A subdomain enumeration tool leveraging diverse techniques, designed for advanced pentesting operations

53
kaipanla-data-parser
kaipanla-data-parser Rainynitesky Python

开盘啦 App 数据抓取与解析工具 - 批量抓取板块/个股数据,Token自动更新,mitmproxy流量拦截

53
MahjongKit
MahjongKit erreurt Python

Riichi Mahjong Kit: (1) Game log crawler (sqlite3, json, bs4); (2) Game log preprocessor; (3) Deterministic algorithms library

52
go-crawler-distributed
go-crawler-distributed golang-collection Go

分布式爬虫项目,本项目支持个性化定制页面解析器二次开发,项目整体采用微服务架构,通过消息队列实现消息的异步发送,使用到的框架包括:redigo, gorm, goquer...

52
TwitterCrawler
TwitterCrawler casolxia Java

抓取twitter数据,可根据时间、话题、用户名等条件抓取数据,twitter爬虫

52
armiarma
armiarma migalabs Go

Armiarma is a Libp2p open-network crawler with a current focus on Ethereum's CL network

52
x12306
x12306 0xHJK Python

12306查票助手,一键查询沿途所有站点,先上车后补票,让你的出行更省心。

52
baidu-chain-dog
baidu-chain-dog CoolAcsi Java

百度莱茨狗爬虫。

51
alipay-crawler
alipay-crawler he426100 PHP

支付宝账单爬虫

51
scrapy.dart
scrapy.dart sachaarbonel Dart

Scrapy, a fast high-level web crawling & scraping framework for dart and Flutter

51
SearchX
SearchX LanyuanXiaoyao-Studio Vue

基于规则的跨平台一站式聚合搜索工具

51
usetube
usetube valerebron TypeScript

search & get datas from youtube no google account needed

51
NeedFree
NeedFree InJeCTrL Python

Crawl 100%-discount games on steam

51
facebook-messenger-bot-tutorial
facebook-messenger-bot-tutorial twtrubiks Python

facebook-messenger-bot-tutorial use Python Django

50
GPlayCrawler
GPlayCrawler KopLyf Python
50
Timbr_V1
Timbr_V1 lvyachao JavaScript

A web service that turns an arbitrary web page into structural JSON data and easy-to-use APIs with just a few clicks

50
html-query
html-query h12w Go

A fluent and functional approach to querying HTML

50
bloodhound
bloodhound vitorfs Python
50
nasty
nasty lschmelzeisen Python

NASTY Advanced Search Tweet Yielder

50
Mini-Spider
Mini-Spider zhangyunhao116 Python

简单、实用的爬虫工具,仅需四步创建属于你的爬虫程序!

50
instagram_scraper
instagram_scraper jbinfo

Extract instagram users informations from hashtags. This scraper can extract emails addresses from Bio section and business email.

50
kepub
kepub TerakomariGandesblood C++

Crawl novels from sfacg, ciweimao, esjzone, lightnovel and masiro; generate, append and extract epub

50
nextcrawler
nextcrawler g089h515r806 JavaScript

Next Crawler 是使用Playwright + Next.js + Prisma等主流技术搭建的网页数据采集器,通过可视化的UI进行配置,即可周期性的通过Playwright驱动浏览器爬取网页数...

50