Topic

scraping

Repositories (1849)

go-crawler
go-crawler lizongying Go

A web crawling framework implemented in Golang, it is simple to write and delivers powerful performance. It comes with a wide range of practical middl...

150
WebReaper
WebReaper alex-on-ai C#

AI-native web scraper. Single binary with a bundled Claude Code skill. MIT-licensed alternative to Firecrawl.

149
GoodreadsScraper
GoodreadsScraper havanagrawal Python

Scrape data from Goodreads using Scrapy and Selenium :books:

148
jazz
jazz jazzdotdev Rust

The Scripting Engine that Combines Speed, Safety, and Simplicity

147
fansly-scraper
fansly-scraper agnosto Go

An all-in-one scraper/downloader for Fansly written in go with the aid of A.I. Download content, Record lives, and Interact with post from your favori...

147
VPSLab-Free-Proxy-List
VPSLab-Free-Proxy-List VPSLabCloud

Free proxy list updated every 15 minutes | HTTP, HTTPS, SOCKS4, SOCKS5, elite & anonymous proxies. Powered by VPSLab.

146
sasori
sasori karthikuj JavaScript

Sasori is a dynamic web crawler powered by Puppeteer, designed for lightning-fast endpoint discovery.

146
cinemaempoa
cinemaempoa cumbucadev Python

Site que agrega filmes em cartaz em algumas das diversas salas de cinema de Porto Alegre.

146
spiderbuf
spiderbuf hhuayuan Python

Spiderbuf 是一个专注于 Python 爬虫练习的网站。提供丰富的爬虫教程、爬虫案例解析和爬虫练习题。Python爬虫开发强化练习,在矛与盾的攻防中不断提高技术水平,...

145
sqrape
sqrape cathalgarvey Go

Simple Query Scraping with CSS and Go Reflection (MOVED to Gitlab)

143
arxiv-miner
arxiv-miner valayDave Python

arxiv_miner is a toolkit for mining research papers on CS ArXiv.

143
abx-dl
abx-dl ArchiveBox Python

⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless...

143
od-database
od-database simon987 Python

Distributed crawler, database and web frontend for public directories indexing

142
html2rss
html2rss html2rss Ruby

📰 Build RSS 2.0 feeds from websites (and JSON APIs) automatically or with a few CSS selectors.

141
nimquery
nimquery GULPF Nim

Nim library for querying HTML using CSS-selectors (like JavaScripts document.querySelector)

140
double-agent
double-agent unblocked-web TypeScript

A test suite of common scraper detection techniques. See how detectable your scraper stack is.

138
wget-lua
wget-lua ArchiveTeam C

Wget-AT is a modern Wget with Lua hooks, Zstandard (+dictionary) WARC compression and URL-agnostic deduplication.

138
lambda-scraper
lambda-scraper teticio JavaScript

Use AWS Lambda functions as a proxy pool to scrape web pages.

137
curl-post-requests
curl-post-requests oxylabs

Learn how to send POST requests with cURL.

136
TelegramAdderTool
TelegramAdderTool saifalisew1508 Python

An Telegram Mass Members Adding/Scraping Tool Written In Python Using Pyrogram Library.

136
linkedin-easyapply-using-AI
linkedin-easyapply-using-AI srikar-kodakandla Python

Automate your LinkedIn job applications with AI! This bot utilizes GPT models such as GPT-4, GPT-3.5, and Google's Gemini Pro for Easy Apply form fill...

134
justetf-scraping
justetf-scraping druzsan Python

Scraping the justETF

134
Web-Data-Scraper
Web-Data-Scraper umbrellaDocumentation JavaScript

Web Data Scraper - no-code internet scraping. Extract and export to CSV, Excel, JSON, Google Sheets, and Webhook.

133
ctenopharyngodon-idella
ctenopharyngodon-idella touero Java

Use the MapReduce's Java interface to distributed crawle the data of Chinese universities and learn basic knowledge of hdfs.

133
pastepwn
pastepwn d-Rickyy-b Python

Python framework to scrape Pastebin pastes and analyze them

130
ScrapeMate
ScrapeMate hermit-crab JavaScript

Scraping assistant tool. Editing and maintaining CSS/XPath selectors across webpages.

130
MachineLearning
MachineLearning yug95 Jupyter Notebook

Machine learning for beginner(Data Science enthusiast)

130
nintendeals
nintendeals fedecalendino Python

Library with a set of tools for scraping information about Nintendo games and its prices across all regions (NA, EU and JP).

129
seleniumcrawler
seleniumcrawler voliveirajr Python

An example using Selenium webdrivers for python and Scrapy framework to create a web scraper to crawl an ASP site

128
Instagram-to-discord
Instagram-to-discord fernandod1 Python

Monitor instagram user account and automatically post new images to discord channel via a webhook. Working 2022!

127
open-australian-legal-corpus-creator
open-australian-legal-corpus-creator isaacus-dev Python

The code used to create and update the Open Australian Legal Corpus, the first and only multijurisdictional open corpus of Australian legislative and...

127
goClone
goClone shurco Go

🌱 goClone - clone websites in seconds

127
robox
robox danclaudiupop Python

Simple library for exploring/scraping the web or testing a website you’re developing

127
FCaptcha
FCaptcha WebDecoy JavaScript

Detect bots, vision AI agents, and headless browsers through 40+ behavioral signals and SHA-256 proof of work. Self-hosted, privacy-first, and fully o...

126
htmlSQL
htmlSQL hxseven PHP

htmlSQL is a experimental PHP library which allows you to access HTML values by an SQL like syntax.

124
viewstate
viewstate yuvadm Python

ASP.NET View State Decoder

123
scout-lang
scout-lang maxmindlin Rust

A web crawling programming language

122
automated-web-scraper-autoscraper
automated-web-scraper-autoscraper oxylabs

This tutorial shows how to automate your web scraping processes using AutoScaper – one of Python web scraping libraries available.

122
shotstars
shotstars snooppr Python

An advanced tool for checking GitHub repositories, with star statistics, including fake star analysis and data visualization.

119
bots-zoo
bots-zoo antoinevastel JavaScript
118
EzSolver
EzSolver ismoiloffS Python

Cloudflare Turnstile solver & bypass — Python, real Chrome browser, no paid APIs. Local HTTP API service included. Auto-solves invisible and managed (...

117
rust-scraping
rust-scraping itehax Rust

Web scraping using rust !

115
media-search-engine
media-search-engine conflict-investigations Python

Search geolocations for (social) media posts in databases like Bellingcat, Cen4InfoRes etc.

115
scrapegraph-mcp
scrapegraph-mcp ScrapeGraphAI Python

ScapeGraph MCP Server

114
qcrawl
qcrawl crawlcore Python

qcrawl - fast async web crawling & scraping framework for Python.

114
scraper
scraper get-set-fetch TypeScript

Nodejs web scraper. Contains a command line, docker container, terraform module and ansible roles for distributed cloud scraping. Supported databases:...

113
rs-bed-covid-indo-api
rs-bed-covid-indo-api satyawikananda TypeScript

API ketersediaan rumah sakit dan tempat tidur rumah sakit untuk pasien covid-19 ataupun non-covid yang berada di Indonesia

110
scrapy-puppeteer
scrapy-puppeteer clemfromspace Python

Scrapy + Puppeteer

110
awesome-ai-web-scraping
awesome-ai-web-scraping h4ckf0r0day

A curated list of AI-powered web scraping tools, LLM-friendly crawlers, MCP servers, and infrastructure for turning the web into data.

110
CC_Scrapper
CC_Scrapper AngelSecurityTeam Python

Telegram CC Scrapper - Debit/Credit Card [channel public or private / group ]

109