微博数据采集,后续会加上知乎,贴吧,小红书,抖音,快手等主流媒体内容
A web crawler (for bug hunting) that gathers more than you can imagine.
pylinkvalidator is a standalone and pure python link validator and crawler that traverses a web site and reports errors (e.g., 500 and 404 errors) enc...
Scrape data from Goodreads using Scrapy and Selenium :books:
A simple library to scrape Pinterest images.
A lite distributed Java spider framework :-)
Web Data Scraper - no-code internet scraping. Extract and export to CSV, Excel, JSON, Google Sheets, and Webhook.
Ruby gem to detect bots and crawlers via the user agent
一些爬虫的代码
✨ 萌萌哒的 AI 网页数据提取助手 ✨
保存百度贴吧帖子到本地,并且支持图片, 视频, 语音等内容。与本项目配套的阅读器 TiebaReader(https://github.com/Sorceresssis/TiebaReader)
Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.
B站用户爬虫 好耶~是爬虫
🗿 npm ↔️ Algolia replication tool :skier: :snail: :artificial_satellite:
🤖Free Agent Line Bot with Web Search, Google Image Search, Image Generator, Video Generator...
AI-native web scraper. Single binary with a bundled Claude Code skill. MIT-licensed alternative to Firecrawl.
功能齐全的Pixiv第三方客户端 免代理 支持查看动图查看小说
A utility package for automating lighthouse reporting
宁稳网(旧富投网)、集思录可转债数据&策略分析
A new generation of multi-process async event-driven spider engine based on workerman. Support headless browser. 🌿基于workerman实现的多进程异步事件...
POOPAK - TOR Hidden Service Crawler
A self-hosted, extensible manga reader and download tool with plug-in support.
This is a course-downloader to help NTU students download courses data from NTU Ceiba.
《数据采集从入门到放弃》源码。内容简介:爬虫介绍、就业情况、爬虫工程师面试题 ;HTTP协议介绍; Requests使用 ;解析器Xpath介绍; MongoDB与MySQL; 多线程...
Take a snapshot of any website.
Price tracker monitors of products and alerts you when prices drop. Supported tiki.vn, shopee, lotte.vn, ... Built with firebase https://pricetrack.we...
微博爬虫,一个基于Scrapy框架的轻量微博爬虫,Sina Weibo Spider
Parse through any sitemap in Node.js
Wget-AT is a modern Wget with Lua hooks, Zstandard (+dictionary) WARC compression and URL-agnostic deduplication.
Download books from bookwalker.jp/bookwalker.com.tw
scrapy专利爬虫(停止维护)
Updated lists of IP addresses/whitelists of good bots and crawlers. Includes GoogleBot, BingBot, DuckDuckBot, etc.
A php crawler that finds emails on the internets
Grabs all of the audio files from all of the Blinkist books
Spiderbuf 是一个专注于 Python 爬虫练习的网站。提供丰富的爬虫教程、爬虫案例解析和爬虫练习题。Python爬虫开发强化练习,在矛与盾的攻防中不断提高技术水平,...
Python ProxyPool for web spider.ProxyPool 是一个用于采集、验证和管理代理IP的轻量级工具,旨在帮助用户自动维护高质量的代理池,方便在爬虫、网络请求中灵活...
Parser and database to index the terpene profile of different strains of Cannabis from online databases
Leetcode Contest Ranking Searcher
哔咔漫画收藏夹下载程序
轻量、异步、开箱即用的社交媒体聚合解析库
爬虫管理系统,支持集群,弹性伸缩。支持运行feapder、scrapy、selenium、playwright等各种框架及脚本
A news crawler for BBC News, Reuters and New York Times.
A powerful MCP server extension providing web search and content extraction capabilities. Integrates DuckDuckGo search functionality and URL content e...
Multithreading download all HD photos / pictures from someone's Sina Weibo album.
Attempts to crawl the Ethereum network of valid Ethereum execution nodes and visualizes them in a nice web dashboard.
Multi-threaded web scraper to download all the tutorials from www.learncpp.com and convert them to PDF files concurrently.
Scraply a simple dom scraper to fetch information from any html based website
SimFin's open source PDF crawler
This repository is no longer maintained.
大麦抢票脚本案例