Topic

scraping

Repositories (1837)

linkedin-scraper
linkedin-scraper akramaznakour JavaScript

Enhanced LinkedIn Job Search Chrome Extension

39
tvseries
tvseries athityakumar HTML

TV Series is a tool that scrapes Episode Synopsis' of popular TV Series' from websites like Wikipedia / IMDb and show in one place with a user-friendl...

38
fulldom-server
fulldom-server strugee JavaScript

Proxy-like server that will show you the DOM of a page after JS runs

38
extract-social-media
extract-social-media fluquid Python

Extract social media links and account names from websites.

38
dergipark-skill
dergipark-skill saidsurucu JavaScript

DergiPark academic search as a Claude-in-Chrome skill (keyword + advanced field search, PDF→text, references) — no CAPTCHA, runs in your own browser.

38
Whatsapp-Scraper
Whatsapp-Scraper In-vincible Python

Scraps all the open chats, and their last n messages, and saves them in a csv file

38
Spy.pet-Info
Spy.pet-Info ThatSINEWAVE JavaScript

This repository serves as an index for all info the community has gathered on the Spy.pet situation and as well as my own tables and tools written for...

38
ETF-Scraper
ETF-Scraper nikulpatel3141 Python

Scrape public ETF and Mutual Fund holdings information

38
lInkedIn-reverese-lookup
lInkedIn-reverese-lookup harsha-iiiv JavaScript

🔎Search LinkedIn profile by email address📧

38
Python-scraper-tutorial
Python-scraper-tutorial Decodo Python

A short introduction to scraping with Python with given steps and an example scraper script.

38
Kemono-Scraper
Kemono-Scraper 3dnsfw TypeScript

UPDATED 2026: Kemono / Coomer Scraper / Downloader / Auto-Retry / Compression

38
scrapeops-scrapy-sdk
scrapeops-scrapy-sdk ScrapeOps Python

Scrapy extension that gives you all the scraping monitoring, alerting, scheduling, and data validation you will need straight out of the box.

38
papercut
papercut armand1m TypeScript

Papercut is a scraping/crawling library for Node.js built on top of JSDOM. It provides basic selector features together with features like Page Cachin...

38
freesoccer
freesoccer andrelmlins TypeScript

:soccer: Free API with results from national soccer competitions

37
google-scraper
google-scraper samaybhavsar PHP

This class can retrieve search results from Google.

37
policy-data-analyzer
policy-data-analyzer wri-dssg-omdena Jupyter Notebook

Building a model to recognize incentives for landscape restoration in environmental policies from Latin America, the US and India. Bringing NLP to the...

37
contact-use
contact-use browser-use HTML

✉️ Use the power of browser-use to contact any person or organization... by any means necessary

37
sneakpeek
sneakpeek flulemon Python

Sneakpeek is a framework that helps to quickly and conviniently develop scrapers. It’s the best choice for scrapers that have some specific complex sc...

37
mangareader-api
mangareader-api stabldev Python

A Python based web scraping api built with fastapi to get manga contents.

37
n8n-ai-instagram-scraper
n8n-ai-instagram-scraper Peter-SB Python

Self hosted AI workflow for scraping Instagram Reels (audio and description). Extracting, summarising and categorising, then storing all relevant info...

37
substack_scraper
substack_scraper bytewife Rust

A scraper for Substack article text content

37
scrapingai
scrapingai Agenty TypeScript

Build web scraping agents using AI to auto-extract the data from websites, capture screenshot, generate pdf from URL and web crawling with Agenty

37
chirps
chirps schedutron Python

Twitter bot powering @arichduvet

36
webradio-metadata
webradio-metadata adblockradio JavaScript

Collection of scraping recipes to get metadata about what is being streamed on webradios

36
KOLscan-leaderboard-scraping
KOLscan-leaderboard-scraping marksantiago290 Python

Python-based web scraper that extracts cryptocurrency trader performance data from KOLscan's leaderboard.

36
poketo
poketo poketo JavaScript

Node library for scraping manga sites

36
api-flight.com
api-flight.com fgparamio HTML

Main API Flight Git Repository

36
tripadvisor-scraper
tripadvisor-scraper andorfermichael Python

Scrape the hotel reviews of a whole city on TripAdvisor

36
PastebinScrapy
PastebinScrapy apurvsinghgautam Python

Threat hunting tool for scraping latest scrapes from Pastebin

36
markupever
markupever awolverp Rust

The fast, most optimal, and correct HTML & XML parsing library for Python written in Rust.

36
botasaurus-starter
botasaurus-starter omkarcloud TypeScript

🚀 OFFICIAL STARTER TEMPLATE FOR BOTASAURUS SCRAPING FRAMEWORK 🤖

36
facebook-discussion-tk
facebook-discussion-tk internaut Python

A collection of tools to (semi-)automatically collect and analyze data from online discussions on Facebook groups and pages.

35
InstaBot
InstaBot drbuche Python

Simple and friendly Bot for Instagram, using Selenium and Scrapy with Python.

35
subio-mcp
subio-mcp alijancb TypeScript

MCP server that scrapes public posts from X, LinkedIn and Hacker News in a local browser that is never signed in

35
jmd_imagescraper
jmd_imagescraper joedockrill Jupyter Notebook

Image scraping library for creating deep learning datasets

35
SneakerBot
SneakerBot mridulghanshala Python

Buy limited edition sneakers

35
rebrowser-puppeteer-core
rebrowser-puppeteer-core rebrowser TypeScript

A drop-in replacement for puppeteer-core patched with rebrowser-patches. It allows to pass modern automation detection tests.

35
geetest-captcha-solver
geetest-captcha-solver ScraperBox-Github JavaScript

Solve the Geetest slider captcha with Puppeteer

35
chromedl
chromedl rusq Go

Go library for scraping or downloading files bypassing Cloudflare protection and browser checks

35
serritor
serritor peterbencze Java

Serritor is an open source web crawler framework built upon Selenium and written in Java. It can be used to crawl dynamic web pages that require JavaS...

34
LeadGen
LeadGen FujiwaraChoki Python

This program uses a GoLang Google Maps Scraper to scrape Google Maps for places and then scrapes the email for each place based on their website.

34
headless-task-server
headless-task-server luka-dev TypeScript

A headless browser task/job queue & runner based on Hero (Chrome)

34
timetable-grabber-sit
timetable-grabber-sit JustBrandonLim TypeScript

Timetable Grabber - SIT is a tool that allows you to grab and export your trimester's timetable to the .ics format where you can import it to your fav...

34
ProductHunt-scraper
ProductHunt-scraper fernandod1 Python

Producthunt.com famous website scraper script. Scrap all offers and save in spreadsheet excel file.

34
ted-scraper
ted-scraper corralm Python

🎙️ TED Talks web scraper

34
BSCWhalesMonitor
BSCWhalesMonitor alessio-ds Python

BSC whales monitor and detector written in Python.

34
python-web-scrapping
python-web-scrapping sallamy2580 Python

Detailed web scraping tutorials for dummies with financial data crawlers on Reddit WallStreetBets, CME (both options and futures), US Treasury, CFTC,...

34
proxi
proxi nicksherron Go

Proxy pool. Finds and checks proxies with rest api for querying results. Can find over 25k proxies in under 5 minutes.

33
deepstate-map-data
deepstate-map-data cyterat Jupyter Notebook

DeepState Map | Occupied | GeoJSON Multipolygon | Daily update

33
bbb
bbb Fedia Svelte

Browser Bot Bookmarklet

33