Topic

scraping

Repositories (1837)

local-api-client-typescript
local-api-client-typescript kameleo-io TypeScript

Official JavaScript/TypeScript library for interacting with Kameleo Client

45
activesoup
activesoup jelford Python

A headless pure-python browser for the web

44
bluebird
bluebird labteral Python

Unofficial Python client for Twitter

44
acon
acon WillyEverGreen Python

The intelligence layer for any web scraper. Pair with Scrapling, Playwright, or httpx to crawl smarter.

44
xdsl-exporter
xdsl-exporter Dentrax Go

xDSL Prometheus Exporter

44
scraperecon
scraperecon DaKheera47 Python

CLI recon tool for scraper developers. Detects TLS fingerprinting, JS challenges, bot protection, and rate limits across 4 stages

44
Extracty
Extracty Mamdouh66 Python

Extract structured data from any unstructured web page

44
Rotating-Proxies-With-Python
Rotating-Proxies-With-Python oxylabs Python

Learn about how to rotate proxies by using Python.

44
fake-http-header
fake-http-header MichaelTatarski Python

A python package to generate random request fields for a http header.

44
scrape-github-trending
scrape-github-trending transitive-bullshit JavaScript

Tutorial for web scraping / crawling with Node.js.

43
RARBG-scraper
RARBG-scraper evyatarmeged Python

With Selenium headless browsing and CAPTCHA solving

43
OSINTLAB
OSINTLAB Purpl3-Dev Shell

This script automates the installation of 50 OSINT tools for reconnaissance and information gathering.

43
super-scraper
super-scraper apify TypeScript

Generic REST API for scraping websites. Drop-in replacement for ScrapingBee, ScrapingAnt, and ScraperAPI services. And it is open-source!

43
shup
shup pystardust Shell

A POSIX shell script to parse HTML

43
scrapingant-client-python
scrapingant-client-python ScrapingAnt Python

ScrapingAnt API client for Python.

43
scrapy-zyte-api
scrapy-zyte-api scrapy-plugins Python

Zyte API integration for Scrapy

43
gwaripper
gwaripper nilfoer Python

Tool for conveniently downloading audios from r/gonewildaudio and similar subreddits

43
laravel-scrapingbee
laravel-scrapingbee ziming PHP

PHP Laravel Library for Scrapingbee Web Scraping API. AI querying supported. Also support Google, Walmart, Amazon, YouTube scraping

43
webmagician-ui
webmagician-ui Jkanon TypeScript

An admin UI project for a configurable web crawler platform

42
Architeuthis
Architeuthis simon987 Go

MITM HTTP(S) proxy with integrated load-balancing, rate-limiting and error handling. Built for automated web scraping.

42
israeli-supermarket-scarpers
israeli-supermarket-scarpers OpenIsraeliSupermarkets Jupyter Notebook

A python package with client to scrape the israeli supermarkets data

42
TorScrapper
TorScrapper little-endian-0x01 Python

A Scraper made 100% in Python using BeautifulSoup and Tor. It can be used to scrape both normal and onion links. Happy Scraping :)

42
chew
chew mmatongo Go

Chew is a Go library for processing various content types into markdown/plaintext.

42
reverseloom
reverseloom KuiChi-x Python

Open-source AI browser agent for web reverse engineering. Turn websites into browser-free API clients and crawlers with CDP, network tracing, and Java...

42
noscrape
noscrape schoenbergerb TypeScript

This repository is deprecated

42
goGetJS
goGetJS davemolk Go

a tool for extracting, searching, and saving JavaScript files (with optional headless browser)

42
UFC-DataLab
UFC-DataLab komaksym Jupyter Notebook

UFC Fights Dataset Collection and Analysis

42
crawlee-cloud
crawlee-cloud crawlee-cloud TypeScript

Self-hosted, open-source platform for running Apify Actors. Drop-in compatible with the Apify SDK.

42
TikDown
TikDown xtekky Python

Fast TikTok NO Watermark Video Downloader (username or url)

42
movie-posters-convnet
movie-posters-convnet adrz Python

Unsupervised clustering of movie posters with features extracted from Convolutional Neural Network

41
hyper-sdk-playwright
hyper-sdk-playwright Hyper-Solutions TypeScript

Hyper Solutions SDK for Playwright - Bypass Akamai Bot Manager, Incapsula, Datadome and Kasada.

41
webtranspose
webtranspose mike-gee Python

Web scraping API for building AI applications.

41
lc-webscraping
lc-webscraping carpentries-incubator Python

Introduction to web scraping

41
crawlee-one
crawlee-one JuroOravec TypeScript

Production-ready web scraping in a single function call. Built on Crawlee.

41
raiplay-dl
raiplay-dl wetcork Python

The most advanced raiplay.it downloader

41
html-table-to-json
html-table-to-json bndnsmth JavaScript

Generate JSON representations of HTML tables

40
gopher-parse-sitemap
gopher-parse-sitemap oxffaa Go

A high effective golang library for parsing big-sized sitemaps and avoiding high memory usage. The sitemap parser was written on golang without extern...

40
nhasixapp
nhasixapp shirokun20 Dart

Unofficial NHentai mobile app with flutter and bloc

40
TikTok-Live-Api
TikTok-Live-Api EulerStream C#

TikTok LIVE API Client - The #1 Worldwide freemium SaaS API for TikTok LIVE data retrieval, offered in all major languages!

40
pyplexity
pyplexity citiususc Python

Cleaning tool for web scraped text

40
linkeBot
linkeBot fabiodeandrade HTML

🔎 um bot de Web Scraping para mostrar vagas do LinkedIn

40
rebrowser-puppeteer
rebrowser-puppeteer rebrowser

A drop-in replacement for puppeteer patched with rebrowser-patches. It allows to pass modern automation detection tests.

40
vegapull
vegapull coko7 Rust

👒 One Piece TCG data scraper written in Rust

40
myanimelist-data-set-creator
myanimelist-data-set-creator debakarr Python

Collection of some simple python scripts to create https://myanimelist.net/ anime and user data set.

39
cloudflare-bypass
cloudflare-bypass HasData JavaScript

This repository provides minimal working examples for bypassing Cloudflare 1020 errors using Playwright in both Python and Node.js. The focus is on sh...

39
everand-downloader
everand-downloader CrazyCoder76 Python

This is everand book downloader

39
etf4u
etf4u leoncvlt Python

📊 Python tool to scrape real-time information about ETFs from the web and mixing them together by proportionally distributing their assets allocation

39
CobWeb-lnx
CobWeb-lnx GoncaloMark Python

CobWeb is a Python library for web scraping. The library consists of two classes: Spider and Scraper.

39
zoominfo_scraper
zoominfo_scraper ScrapingAnt Python

Zoominfo scraper with using of rotating proxies and headless Chrome from ScrapingAnt

39
linkedin-scraper
linkedin-scraper akramaznakour JavaScript

Enhanced LinkedIn Job Search Chrome Extension

39