Topic

crawler

Repositories (1456)

Chemrtron
Chemrtron cho45 TypeScript

Chemr is a document viewer; fuzzy match incremental search.

64
crawdad
crawdad schollz Go

Cross-platform persistent and distributed web crawler :crab:

64
medium-crawler
medium-crawler NISH1001 Python

A crawler for scraping posts from medium.com

64
Auto_Shadowsocks
Auto_Shadowsocks VonSdite Python

Shadowsocks. 科学上网, 仅供学习。是免费的服务器,可能存在科学上网不稳定。

63
SoFIFA
SoFIFA DiogoDantas Jupyter Notebook

A SoFIFA webcrawler and Machine Learning prediction

63
bthello
bthello rehe0x Python

Python3 DHT 磁力种子爬虫 种子解析 种子搜索 演示地址

63
ZhihuVAPI
ZhihuVAPI cheezone Python

优雅地玩知乎

62
koshort
koshort koshort Python

(deprecated) :cat: koshort is a Python package for Korean internet spoken language crawling and processing... or maybe Korean domestic cat.

62
sciBASIC
sciBASIC xieguigang Visual Basic .NET

sciBASIC# is a kind of dialect language which is derive from the native VB.NET language, and written for the data scientist.

62
Java-Carwler-Technology
Java-Carwler-Technology soberqian Java

网络数据采集技术—Java网络爬虫 (书稿完整代码,涉及网络爬虫的各种技术和知识点)

62
m3u8Downloader
m3u8Downloader mrzhangfelix Python

meijuba.net,Python crawler,M3U8格式视频下载,桌面应用

62
crawler-project
crawler-project Albert-W Go

Google资深工程师深度讲解Go语言 爬虫项目。

62
crawler_JD_what_worthy_buying
crawler_JD_what_worthy_buying HarborZeng Python

爬取京东商品所有评论,利用情感分析,判断商品是否值得买

62
scrapy-distributed
scrapy-distributed Insutanto Python

A series of distributed components for Scrapy. Including RabbitMQ-based components, Kafka-based components, and RedisBloom-based components for Scrapy...

62
novel-downloader
novel-downloader yjqiang Python

万能小说下载器

62
facebook-page-info-scraper
facebook-page-info-scraper wael-sudo2 Python

Free Facebook pages MetaData Scraping Library - Unlimited Calls

62
Tapestry
Tapestry NatsuFox Python

Tapestry - 基于 Agent Skill Bundle 的轻量级书签知识库

62
aio-vextractor
aio-vextractor panoslin Python

解析视频 网站/APP/H5 页面视频信息。支持抖音、腾讯视频、YouTube、Instagram 等40余个网站与APP

61
a11y-sitechecker
a11y-sitechecker forsti0506 TypeScript

Automatic accessibility checker with website crawling + screenshots for easy use

61
webrtc-local-ip-leak
webrtc-local-ip-leak niespodd HTML

Oh no, stop this. You can see my local IP address 😲! Use `foundation` attribute against CRC32 lookup table to reveal local IP address of a Chrome/Chr...

61
mcp
mcp supadata-ai TypeScript

Official Supadata MCP Server - Adds powerful video & web scraping to Cursor, Claude and any other LLM clients.

61
TumblTwo
TumblTwo johanneszab C#

TumblTwo, an Improved Fork of TumblOne, a Tumblr Downloader.

60
WebCrawler
WebCrawler Misterhex C#

Just a simple web crawler which return crawled links as IObservable using reactive extension and async await.

60
WebSpider
WebSpider xdoer JavaScript

基于Nodejs,superagent,cheerio的在线web爬虫项目,支持生成API

60
findopendata
findopendata findopendata Python

A search engine for Open Data

60
rolling-news
rolling-news Jacen789 Python

获取滚动新闻

60
snapcrawl
snapcrawl DannyBen Ruby

Crawl a website and take screenshots

60
custom-crawler
custom-crawler rollrat C#

🌌 High productivity semi-automatic crawler generator 🛠️🧰

60
crawler-userscript
crawler-userscript zjh1943 JavaScript

一个基于 Tampermonkey 插件平台开发的爬虫。主要目的是最大限度模拟用户环境,避免被反爬虫系统识破。

60
DDoM
DDoM Endermanch Python

A simple, open-source, easy to use, and free download manager for malware samples.

60
thecrowler
thecrowler pzaino Go

A Content Discovery and Development Platform. Empowering Cybersecurity, AI, Marketing, and Finance professionals and researchers to discover, analyze,...

60
pomp
pomp estin Python

Screen scraping and web crawling framework

59
lyrics-crawler
lyrics-crawler willamesoares Python

Get the lyrics for the song currently playing on Spotify

59
google_news_scraper_and_sentiment_analyzer
google_news_scraper_and_sentiment_analyzer pratikpv Python

Downloads news articles from Google news and uses pre-trained NLP models to perform sentiment analysis

59
Web-Iota
Web-Iota SatinWukerORIG Python

Iota is a web scraper that can find all of the images and links/suburls on a webpage

59
billboard-json
billboard-json KoreanThinker TypeScript

🎧 Get json type billboard hot 100 chart

59
spider-nodejs
spider-nodejs spider-rs Rust

Spider ported to Node.js

59
python
python joaopauloaramuni HTML

Repo Python

59
phpcrawl
phpcrawl mmerian PHP

Copy of http://phpcrawl.cuab.de/ for using with composer

58
Daily-code
Daily-code rui7157 Python

日常代码爬虫、gui小工具等

58
tool-gin
tool-gin bajins Go

基于go-gin框架建立减少冗余动作项目,如:下载一些工具

58
proxycrawl-python
proxycrawl-python crawlbase Python

ProxyCrawl Python library for scraping and crawling

58
gscholar-citations-crawler
gscholar-citations-crawler thu-pacman Python

Crawl all your citations from Google Scholar

58
simple_bank_korea
simple_bank_korea Beomi Python

simple crawler for Korean banks with Transactions

58
unfx-proxy-parser
unfx-proxy-parser openproxyspace JavaScript

Unfx Proxy Parser - Nextgen proxy parser with deep links crawler. Follow to internal links, third-party links. Sorting results by countries.

58
site-md
site-md yazinsai TypeScript

Serve clean Markdown from your Next.js site to AI agents, crawlers, and LLMs. Humans get HTML, agents get clean Markdown of the same pages. Two-file i...

58
talospider
talospider howie6879 Python

talospider - A simple,lightweight scraping micro-framework

57
All-IT-eBooks-Spider
All-IT-eBooks-Spider Kulbear Python

[Updated] A simple python crawler for my tutorial blog at http://www.jianshu.com/p/8fb5bc33c78e

57
instagram-hashtag-crawler
instagram-hashtag-crawler simonseo Python

Crawl Instagram hashtags

57
SearchEngineScrapy
SearchEngineScrapy naqushab Python

Scrape data from Google.com, Bing.com, Baidu.com, Ask.com, Yahoo.com, Yandex.com

57