Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining"
API of DouYin for Humans used to Crawl Popular Videos and Musics
NetDiscovery 是一款基于 Vert.x、RxJava 2 等框架实现的通用爬虫框架/中间件。
Python的基础练习代码与各种爬虫代码
Locally saves webpages to your hard disk with images, css, js & links as is.
爬取菜鸟教程网站并转PDF__python_crawer_by_chrome
Moodle-DL downloads course content fast from Moodle (eg. lecture pdfs)
Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change.
带你了解一下Golang的市场行情
What do people have in their dotfiles?
Prying Deep - An OSINT tool to collect intelligence on the dark web.
LinkedIn Scraper (currently working 2020)
Jie stands out as a comprehensive security assessment and exploitation tool meticulously crafted for web applications. Its robust suite of features en...
Free Web Scraping Tool with Java
👩 美女写真套图爬虫(一)
Simple yet powerful automation stuffs.
a reliable high-level web crawling & scraping framework for Node.js.
Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.
Crawljax
swiss army knife for hackers
FreeProxy: Collecting free proxies from internet. (全球海量高质量免费代理,支持爬取数十个免费代理分享源,支持自定义规则代理筛选,爬虫与数据分析必备,...
Fresh Onions is an open source TOR spider / hidden service onion crawler hosted at zlal32teyptf4tvi.onion
Crawl and extract (regular or onion) webpages through TOR network
🌈Python3网络爬虫实战:QQ音乐歌曲、京东商品信息、房天下、破解有道翻译、构建代理池、豆瓣读书、百度图片、破解网易登录、B站模拟扫码登录、小鹅通、荔枝微课
Crawler for Nintendo Switch eShop
Open-source Enterprise Grade Search Engine Software
A framework for creating semi-automatic web content extractors
a new crawler based on python with more function including Network fingerprint search
Html网页正文提取
ScraperAI is an open-source, AI-powered tool designed to simplify web scraping for users of all skill levels.
A very simple news crawler with a funny name
⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper API,使用本地运行的Whisper模型进行推理,并支持多GPU并发,针对分布式部署进行设计...
台灣股票即時爬蟲。Taiwan Stock Exchange Real Time Crawler
澎湃新闻,新浪新闻,腾讯新闻,搜狐新闻,新闻联播,泰晤士报,纽约时报,BBCNews,旨在爬取所有新闻门户网站的新闻,禁止将所得数据商用!
Script that crawls meta data from ICLR OpenReview webpage. Tutorials on installing and using Selenium and ChromeDriver on Ubuntu.
Real-time detection of anti-bot systems, CAPTCHAs & fingerprinting techniques. Identifies Cloudflare, Akamai, DataDome, reCAPTCHA, hCaptcha, Shape Se...
Easily create XML sitemaps for your website.
Coomer| kemono .party or su downloader
🕵️ Python project to crawl for JavaScript files and search for secrets like API keys, authorization tokens, hardcoded credentials, etc.
Scrapes all photos and videos in a web page / Instagram / Twitter / Tumblr / Reddit / pixiv / TikTok
各种App、小程序、网站的请求签名或加密算法。 现已有:自如、小红书、蛋壳公寓、luckin coffee(瑞幸咖啡)、bangkokair(曼谷航空)
This repository contains all the code I use in my YouTube tutorials.
dude uncomplicated data extraction: A simple framework for writing web scrapers using Python decorators
Search emails from a domain through search engines
golang实现的爬虫框架,使用者只需关心页面规则,提供web管理界面。基于colly开发。
Selenium Open Source Search Engine & crawler
:musical_note: 缓存文件转换为 MP3 文件
Second-order subdomain takeover scanner
Android 本地网络小说爬虫,基于jsoup及xpath
Google search results crawler, get google search results that you need