Topic

crawler

Repositories (1456)

FF14AutoSignIn
FF14AutoSignIn renchangjiu Python

FF14 国服官网自动签到脚本

40
lolcrawler
lolcrawler jonaslejon Python

Headless web crawler for bugbounty and penetration-testing/redteaming

40
doogle
doogle safesploitOrg PHP

Doogle is a search engine and web crawler which can search indexed websites and images

40
Raven
Raven Symbolexe Go

Raven is a powerful and customizable web crawler written in Go.

40
wayurls
wayurls alwalxed Go

CLI tool for fetching URLs from Wayback Machine, Common Crawl, and VirusTotal.

40
crawlee-one
crawlee-one JuroOravec TypeScript

Production-ready web scraping in a single function call. Built on Crawlee.

40
acon
acon WillyEverGreen Python

The intelligence layer for any web scraper. Pair with Scrapling, Playwright, or httpx to crawl smarter.

40
AntiCloudFlare
AntiCloudFlare s045pd HTML

对抗cloudflare载入页反爬虫防护(已失效)

39
auto_crawler_ptt_beauty_image
auto_crawler_ptt_beauty_image twtrubiks Python

自動抓取 ptt 表特板圖片

39
crawler
crawler crawlerclub Go

Crawler4U, a general purpose focused crawler

39
fess-crawler
fess-crawler codelibs Java

Web/FileSystem Crawler Library

39
Android-Apps-Downloader
Android-Apps-Downloader harismuneer Python

📱 A utility for downloading Android apps from the Google Play Store and Xiaomi App Store (the Chinese App Store).

39
selenium_facebook_scraper
selenium_facebook_scraper Mhmd-Hisham Python

A simple python3 script used to download a users's friend list from facebook.

39
papercut
papercut armand1m TypeScript

Papercut is a scraping/crawling library for Node.js built on top of JSDOM. It provides basic selector features together with features like Page Cachin...

39
novelsave_sources
novelsave_sources m-haisham Python

A collection of webnovel sources offering varying amounts of scraping capability.

39
CobWeb-lnx
CobWeb-lnx GoncaloMark Python

CobWeb is a Python library for web scraping. The library consists of two classes: Spider and Scraper.

39
proxy-in-a-box
proxy-in-a-box naiba Go

Automatic proxy pool for web scraping - crawls, validates, and rotates proxies with rate limiting and MITM support

39
python-facebook-bot
python-facebook-bot tudoanh Python

Get facebook events from location with Python 3

38
BaiduImageCrawler
BaiduImageCrawler flexwang-zz Python

A multithreaded tool for downloading search results of Baidu image search.

38
integrada.minhabiblioteca.com.br
integrada.minhabiblioteca.com.br tharyckgusmao JavaScript

Download de livros para PDF/EPUB - Integrada.minhabiblioteca / vitalsource

38
generic-seeder
generic-seeder team-exor C++

Generic altcoin DNS seeder. Compatible with virtually any cryptocurrency cloned from bitcoin. Built-in lightweight DNS server ~ Cloudflare DNS support...

38
ProxyScan
ProxyScan Its-Vichy Go

🔎 scan the internet to find "private" proxies.

38
ExHentaiReader
ExHentaiReader AndyHsiehTA HTML

Best manga-viewer on windows for crawling/downloading/browsing exhentai.

38
EH-PDF
EH-PDF Galgamer-org Python

將一個 E-Hentai 畫廊下載並轉換成 PDF,方便在 Kindle 上閱讀 以及在 iPad 上閱讀並作筆記,,,

38
article_crawler
article_crawler tychozzz Python

✨ Article Crawler is a package used to crawl articles with Markdown format from a specific webpage and store them locally in HTML / Markdown formats.

38
geckordp
geckordp jpramosi Python

A client implementation of Firefox DevTools over remote debug protocol in python

38
WebRecon
WebRecon flashnuke Python

A collection of pentesting web scanners

38
xhs-js
xhs-js saifeiLee JavaScript

基于小红书web端的请求封装,JS实现

38
LLM-Web-Crawler
LLM-Web-Crawler buildship-ai TypeScript

Web Scraper and Crawler for LLM Apps and AI Workflows with NoCode / LowCode. Plug and play with your own logic and customize it flexibly and scalably...

38
NodeSpider
NodeSpider Bin-Huang TypeScript

[DEPRECATED] Simple, flexible, delightful web crawler/spider package

37
toxcrawler
toxcrawler JFreegman C

A Tox DHT network crawler

37
Facebooker
Facebooker gpwork4u Python

an unofficial facebook api

37
Mini-Projects
Mini-Projects nazaninsbr Python

A collection of short projects, you could try and implement these as short projects or use them as part of a larger project.

37
Python
Python fmw666 Python

🍋 Python基础、Pygame游戏编程、Python算法与面试题、四种常用的Python Web框架、爬虫、数据可视化、机器学习。一共七个Python大方向!

37
crawlhtmltopdf
crawlhtmltopdf osdodo Python

一个将runoob.com转换为PDF的爬虫

37
xray_pool
xray_pool allanpk716 Go

基于 Xray-core、glider 的代理池工具

37
Data-Collection-Process-for-the-2024-Huashu-Cup-C-Problem
Data-Collection-Process-for-the-2024-Huashu-Cup-C-Problem Diraw Python

华数杯2024C题数据集收集过程

37
instagram-data-scraper
instagram-data-scraper Z786ZA

Instagram Data Scraper analyze profile

37
sneakpeek
sneakpeek flulemon Python

Sneakpeek is a framework that helps to quickly and conviniently develop scrapers. It’s the best choice for scrapers that have some specific complex sc...

37
awesome-digital-preservation
awesome-digital-preservation ruarxive

Awesome list dedicated to digital and data preservation tools, sources, services and so on.

37
maw
maw memvid TypeScript

Crawl any website into a single searchable file. Query it forever, offline.

37
NetEaseCloudMusicCrawler
NetEaseCloudMusicCrawler timelessmemory Java

HttpClient + Jsoup + Queue

36
MMDownloader
MMDownloader occidere Java

마루마루 다운로더 신규 프로젝트

36
gargantua
gargantua andreaskoch Go

The fast website crawler

36
imooc-crawler
imooc-crawler monkeym4ster JavaScript

[Obsolete] imooc web crawler in Node.js(使用 Node.js 编写的慕课网爬虫)

36
golearn
golearn hackfengJam Go

🔥 Golang basics and actual-combat (including: crawler, distributed-systems, data-analysis, redis, etcd, raft, crontab-task)

36
vw-crawler
vw-crawler vector4wang Java

:beetle:简单轻便的Java爬虫框架,只要会一点简单的正则表达式和简单的css选择器就能轻松的采集数据。

36
node-html-crawler
node-html-crawler safonovpro JavaScript

Simple for use node html crawler (spider) of site web pages

36
PyperGrabber
PyperGrabber pykong Python

Fetches PubMed article IDs (PMIDs) from email inbox, then crawls PubMed, Google Scholar and Sci-Hub for respective PDF files.

36
crazyDhtSpider
crazyDhtSpider ixiaofeng PHP

Based on Swoole,a PHP DHT crawler, which have insane productivity(依托于swoole的PHP版本的DHT爬虫,有着奇高的效率)

36