Topic

crawler

Repositories (1456)

tumblr-crawler-cli
tumblr-crawler-cli tzw0745 Python

Tumblr Download Tool with High Speed and Customization. 高性能&高定制化的Tumblr下载工具。

75
crawler_examples
crawler_examples liuslnlp Python

Some classic web crawler projects.一些经典的爬虫

74
newspaperjs
newspaperjs flickz HTML

News extraction and scraping. Article Parsing

74
light-crawler
light-crawler zhang2333 JavaScript

a simplified directed customizable website crawler

74
app-crawler
app-crawler timschneeb Python

Python script that searches GitHub, F-Droid and IzzySoft's F-Droid repo for apps with Shizuku support. Updated daily.

74
python-testing-crawler
python-testing-crawler python-testing-crawler Python

A crawler for automated functional testing of a web application

74
qr-pirate
qr-pirate mzollin Python

crawl QR-codes from search engines and look for bitcoin private keys

74
ComicSpider
ComicSpider QuantumLiu Python

动漫之家漫画站电脑版原图爬虫

73
Instagram-downloader
Instagram-downloader fernandod1 Python

Instagram user's photos and videos downloader. Download all media files from any username. Working 2022!

73
site-mirror-py
site-mirror-py generals-space Python

[码云](https://gitee.com/generals-space/site-mirror-py) 通用爬虫, 仿站工具, 整站下载

73
meta-spy
meta-spy DEENUU1 Python

👾 CLI MetaSpy (Facebook, Instagram) scraper and crawler - instagram account, facebook accounts, pages and search

72
aio-scrapy
aio-scrapy ConlinH Python

Implement scrapy with asyncio

72
carbonbot
carbonbot crypto-crawler Rust

A command line tool based on the crypto-crawler library.

72
Wedge
Wedge LZ0211 JavaScript

可配置的小说下载及电子书生成工具

72
MusicTagger
MusicTagger Mai-icy Python

一个可以补全mp3,flac文件元数据的图形化界面。还可以下载歌词

72
tech-stack-datasets
tech-stack-datasets leadita

Open datasets of companies & websites grouped by technologies they use (CSV & JSON). Discover who uses Shopify, Stripe, Woocommerce, HubSpot, and more...

71
Pinterest-Crawler
Pinterest-Crawler SajjadAemmi Python

Download HD images from pinterest by your favorite keywords

71
crawlerdetect
crawlerdetect x-way Go

Golang module to detect bots and crawlers via the user agent

70
OpenCrawler
OpenCrawler merwin-asm Python

Open Crawler || Open Source Crawler

70
spider
spider jhao104 Python

python crawler spider

70
Crawling-CV-Conference-Papers
Crawling-CV-Conference-Papers seanywang0408 Jupyter Notebook

Crawling CV conference papers with Python.

70
Tor_Spider
Tor_Spider absingh31 Python

Python project to crawl and scrap the lesser known deep web or one can say dark web. Just provide the onion link and get started.

70
ipfs-crawler
ipfs-crawler trudi-group Go

A crawler for the IPFS network, code for our paper (https://arxiv.org/abs/2002.07747). Also holds scripts to evaluate the obtained data and make simil...

70
tiktok-scraper-php
tiktok-scraper-php snuzi PHP

Tiktok (Musically) PHP scraper

70
lrabbit_scrapy
lrabbit_scrapy litter-rabbit Python

a quick start python mutil thread crawl

70
rubium
rubium vifreefly Ruby

Antidetect Headless Chrome Browser for Ruby Web Scraping and Automation

70
ForgeRSS
ForgeRSS tmwgsicp Python

将任意网站转换为 RSS 订阅源、多引擎抓取、反爬突破、支持爬取抖音、快手、小红书、B站、知乎、小宇宙、知识星球

69
BOJ-AutoCommit
BOJ-AutoCommit ISKU Python

When you solve the problem of Baekjoon Online Judge, it automatically commits and pushes to the remote repository.

69
robotstxt
robotstxt ropensci R

R 📦 for parsing and checking robots.txt files 🤖

69
JewelCrawler
JewelCrawler DMinerJackie Java

豆瓣电影爬虫——a crawler which is able to crawl movie detail and short comments, save them to database mysql, also include Sentiment analysis based on...

69
bilibili_comment_crawl
bilibili_comment_crawl lemonmindyes Python

爬取bilibili视频下的评论,最新出品!!!⚠本代码只适用于学习,做其他事情概不负责!!!

69
slime
slime nekolr Java

🍰 A visual crawler management platform

69
lezhin-comics-downloader
lezhin-comics-downloader ImSejin Java

📥 Downloader for lezhin comics

68
SecretScraper
SecretScraper PadishahIII Python

SecretScraper is a web scraper that crawl through target websites, scrape from http response and extract secret information via regular expression. Ru...

68
nest-crawler
nest-crawler saltyshiomix TypeScript

An easiest crawling and scraping module for NestJS

67
AfdianToMarkdown
AfdianToMarkdown PhiFever Go

爱发电爬虫(afdian.com)

67
hproxy
hproxy howie6879 Python

hproxy - Asynchronous IP proxy pool, aims to make getting proxy as convenient as possible.(异步爬虫代理池)

67
bolsa
bolsa gicornachini Python

Biblioteca feita em Python com o objetivo de facilitar o acesso a dados de seus investimentos na bolsa de valores(B3/CEI) através do Portal CEI.

67
zhihu-crawler
zhihu-crawler NightMarcher Python

徒手实现定时爬取知乎,从中发掘有价值的信息,并可视化爬取的数据作网页展示。

67
js_block
js_block webcoding HTML

研究学习各种拦截:反爬虫、拦截ad、防广告注入、斗黄牛等

67
kabegame
kabegame kabegame Rust

Kabegame — An anime image crawler client with pluggable crawlers (from a GitHub plugin repo), wallpaper rotation, local folder sync album. Supports Wi...

67
Taiwan-Stocks
Taiwan-Stocks smalldan1022 Python

台灣上市櫃公司爬蟲,分析盤後股票趨勢以及繪製K線圖、均線圖、三大法人成交量

67
PSGameSpider
PSGameSpider RavelloH JavaScript

自动爬取所有PlayStationStore中的所有游戏信息,包括封面、描述、价格、评分等,生成网页并索引 # # # Automatically crawl all game infos in all playstation...

67
chat-plugin-web-crawler
chat-plugin-web-crawler lobehub HTML

🧩 / 🕸 WebsiteCrawler - This plugin automatically crawls the main content of a specified URL webpage and uses it as context input.

67
paperCrawler
paperCrawler sucv Python

This is a Scrapy-based web-spider. It scrapes papers from TOP conferences and journals.

66
python-crawler
python-crawler ityouknow Python

Python Crawler

66
github-trending
github-trending doforce Python

GitHub trending repositories and developers APIs for real time, powered by crawlers | 通过爬虫获取 GitHub 热门项目和开发者的实时 API

66
datacrawl
datacrawl DataCrawl-AI Python

A simple and easy to use web crawler for Python

65
Visual_MediaCrawler
Visual_MediaCrawler persist-1 Python

可视化爬虫(支持:哔哩哔哩 | 抖音 | 小红书 | 贴吧 | 微博 | 知乎 | 快手),异步、高效、直观地采集国内主流平台的媒体数据的前后端一体项目(Based on "Medi...

65
dht-crawler
dht-crawler hijkzzz Go

A DHT Crawler based on Goroutine

65