Async Python 3.6+ web scraping micro-framework based on asyncio
Google, Naver multiprocess image web crawler (Selenium)
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
python爬虫,目前库存:网易云音乐歌曲爬取,B站视频爬取,知乎问答爬取,壁纸爬取,xvideos视频爬取,有声书爬取,微博爬虫,安居客信息爬取+数据可视化,哔哩...
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
抖音爬虫——采集账号主页、喜欢、收藏、音乐原声、话题、搜索、合集、作品、关注、粉丝等公开数据。
Scrapoxy hides your scraper behind a cloud. It starts a pool of proxies to send your requests. Now, you can crawl without thinking about blacklisting!
Crawl a website and run it through Google lighthouse
ScopeSentry-Cyberspace mapping, subdomain enumeration, port scanning, sensitive information discovery, vulnerability scanning, distributed nodes
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
Elasticsearch File System Crawler (FS Crawler)
小红书数据采集、网站图片、视频资源批量下载工具,颜值超高的数据采集工具(批量下载,视频提取,图片)Telegram:https://t.me/+ZtLSwuIKTo44MDY1
A web privacy measurement framework
An ergonomic, privacy-aware Python HTTP Client
浏览过的精彩逆向文章汇总,值得一看
Collection of patches for puppeteer and playwright to avoid automation detection and leaks. Helps to avoid Cloudflare and DataDome CAPTCHA pages. Easy...
It makes a preview from an URL, grabbing all the information such as title, relevant texts and images.
🤖 Scrape data from HTML websites automatically by just providing examples
K 哥爬虫代码分享,JS 逆向,爬虫进阶。关注公众号:K哥爬虫
Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.
Python爬虫,京东自动登录,在线抢购商品
The Prime Cross Site Request Forgery (CSRF) Audit and Exploitation Toolkit.
The fastest dork scanner written in Go.
📝 quickly crawl the information (e.g. followers, tags etc...) of an instagram profile.
❤️ Fredy - [F]ind [R]eal [E]state [D]amn Eas[y] - Fredy keeps searching for new apartments, houses, and flats in Germany on platforms like ImmoScout24...
收集各种免费的 Python 爬虫项目
Beanbun 是用 PHP 编写的多进程网络爬虫框架,具有良好的开放性、高可扩展性,基于 Workerman。
一个好用的哔哩哔哩漫画下载器,拥有图形界面,支持关键词搜索漫画和二维码登入,黑科技下载未解锁章节,多线程下载,多种保存格式,本地漫画管理,一键检查更新...
基于appium的app自动遍历工具
massive SQL injection vulnerability scanner
🤖 Fake fingerprints to bypass anti-bot systems. Simulate mouse and keyboard operations to make behavior like a real person.
👧 美女写真套图爬虫(二)
:beers: bilibili video (including bangumi) and danmaku downloader | B站视频(含番剧)、弹幕下载器
A Simple Mihomo GUI. 一个简易的 Mihomo 桌面客户端
Easily download all the photos/videos from tumblr blogs. 下载指定的 Tumblr 博客中的图片,视频
📰 Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.
Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/...
Crawly, a high-level web crawling & scraping framework for Elixir.
Write web scrapers in Ruby using a clean, AI-assisted DSL. Kimurai uses AI to figure out where the data lives, then caches the selectors and scrapes w...
Run a high-fidelity browser-based web archiving crawler in a single Docker container
Scalable Python web scraping scripts for +40 popular domains
A tool for pixiv.net. 人人可用的P站爬虫
Google play scraper for Python inspired by <facundoolano/google-play-scraper>
A scalable, mature and versatile web crawler based on Apache Storm
Golang短视频去水印:抖音,皮皮虾,火山,微视,最右,快手,全民小视频,皮皮搞笑,西瓜视频,虎牙,梨视频,acfun,好看视频...
小说下载|小说爬取|起点|笔趣阁|导出Markdown|导出txt|转换epub|广告过滤|自动校对
✌️ Python3 BitTorrent DHT crawler
SpiderSuite releases, wiki and roadmap
A high performance web crawler / scraper in Elixir.