Topic

crawling

Repositories (1389)

ArtStyle-Detector
ArtStyle-Detector Mirtia Python

A project aiming to detect artstyles from images. It queries Wikimedia Commons to collect images for the training set.

5
kafka-ES-DataPrakiraanCuaca
kafka-ES-DataPrakiraanCuaca RomySaputraSihananda Python

Simulasi transmisi data hasil crawling dari DataPrakiraanCuaca menggunakan Python, Kafka, dan Elasticsearch.

5
GooglePlayDatabaseMirror
GooglePlayDatabaseMirror BaseMax PHP

Repository of designing a crawler script to update a mirror database from Google Play on PHP.

5
Spider
Spider enigma522 Python

This asynchronous web crawler is designed for reconnaissance tasks. It crawls a specified URL up to a defined depth, extracting useful information

5
crwlr
crwlr busterc JavaScript

🕷a minimal puppeteer crawler api

5
CountriesSearchEngine
CountriesSearchEngine anshul1004 Python

A search engine built to retrieve geographical information of any country.

5
bot-safe-agents
bot-safe-agents ivan-sincek Python

A library for fetching a list of bot-safe user agents.

5
sitemapr
sitemapr alphaprime-dev Python

sitemapr is a library that generates sitemaps for SPA websites by reading site structures defined in declarative configuration.

5
langchain-advertools
langchain-advertools eliasdabbas Jupyter Notebook

LangChain integration for advertools

5
Facraw-Playwright
Facraw-Playwright ryyos Python

Facebook scraping using playwright

5
estela-entrypoint
estela-entrypoint bitmakerla Python

estela entrypoint for job runner 🕸

5
tider
tider ZLotusRain Python

A fast, simple, extensible and powerful framework for web crawling.

5
kumo
kumo wihlarkop Rust

An async web crawling framework for Rust — Scrapy for Rust.

5
amazon_luwak_coffee_scraper
amazon_luwak_coffee_scraper omar-elmaria Python

This repo contains a Python-based web crawler that scrapes data on Luwak coffee products from amazon.de. It is designed to surpass Amazon's anti-bot m...

5
Text_mining
Text_mining Jimin980921 Jupyter Notebook

텍스트마이닝을 이용한 소비자분석 _네이버쇼핑 리뷰크롤링

5
Fundamentus_scraping
Fundamentus_scraping GuilhermeUchoa Python

Crawler do site Fundamentus.com com o uso do framework scrapy, tanto da aba detalhada como a de resumo. (Todas as infomações)

5
awesome-web-crawler
awesome-web-crawler ilovedevs

List of best web crawlers to extract data from the web. Find web crawling tools for different needs.

5
Airput
Airput Airsequel Haskell

CLI tool for populating Airsequel with data. Includes a crawler for GitHub repos metadata.

5
Migale
Migale cth-latest C#

Migale was born out of a need to extract data quickly and with a very low development cost. This package is not intended to replace complete and struc...

5
naver_webtoon
naver_webtoon hey411 Python

딥러닝과 머신러닝을 활용한 독자 반응 기반 웹툰 데뷔작 성공 예측 모델

5
EPhoto360
EPhoto360 LordDeveloper PHP

Create text effects online , Effects online for free, photo frames, make face photo montages, custom greeting cards, add vintage filters, turn photos...

5
SiteMapperChromeExtension
SiteMapperChromeExtension MatthewMariner JavaScript

Discover and navigate website structure with smart sitemap detection, visual tree view, and export tools for SEO audits

5
Instagram-image-downloader
Instagram-image-downloader KAispread Kotlin

💟 Instagram Image Downloader

5
go-web-scrapper
go-web-scrapper zahidhasann88 Go

A web scraping API using Golang with Gin and ChromeDP for dynamic site scraping.

5
webArchive
webArchive mrrfv JavaScript

Crawls websites and saves found URLs to a file.

5
craw-kompas
craw-kompas RomySaputraSihananda Python

crawling and scrapping data from the kompas news website

5
RoboNope-nginx
RoboNope-nginx raminf C

Take control of your own content. Enforce access to disallowed web URLs.

5
coronaflight-hkg
coronaflight-hkg poyea JavaScript

😷 Crawler and history manager for dangerous, coronavirus-infected flights to Hong Kong (VHHH)

5
proxycrawl-java
proxycrawl-java crawlbase Java

ProxyCrawl Java library for scraping and crawling

5
Slic
Slic sw-song Python

Single line image classifier

5
Scout
Scout Hemachandra9899 TypeScript

An evidence-first AI research engine with scoped memory, graph intelligence, and agent execution.

5
BiLSTM-StockPrediction-Algorithm
BiLSTM-StockPrediction-Algorithm paulms77 Jupyter Notebook

양방향 LSTM 기반 주가 예측 알고리즘 논문 연구 코드입니다.

4
web-crawler
web-crawler asanakoy C++

Web-Crawler for simple.wikipedia.org on C++

4
laravel-crawler
laravel-crawler ahmedmohamed1101140 PHP

use the app to scrap the product amount from souq amazon or jumia login and give it a try

4
Emotion-Regconition-Youtube
Emotion-Regconition-Youtube ToVinhKhang Jupyter Notebook

Emotion Recognition for Vietnamese Social Media Text (Youtube Comments)

4
Naver-dictionary-crawler
Naver-dictionary-crawler entiff Jupyter Notebook

Crawling Naver dictionary example

4
mediawikiextractor
mediawikiextractor chenmozhijin Python

一个用于从 MediaWiki 网站中提取数据并保存为json的 Python 脚本。|A Python script for extracting data from a MediaWiki website and saving it as json.

4
johnny-cache
johnny-cache Sonictherocketman Python

A simple forward caching proxy. Useful for reducing the bandwidth of polling or crawling public sites.

4
lebonscrap
lebonscrap wbwlkr Python

LeBonScrap is a spider which collect data from Leboncoin.fr, crawl all the pagination links to scrap every ads of the list from one search result of t...

4
SimplePyCrawler
SimplePyCrawler DanielGunna Python

A simple web crawler developed as coursework for Algorithms on Graph Theory - PUC Minas

4
Rotakka
Rotakka Miroka96 TeX

Rotakka is a distributed Akka cluster application designed for scalable Twitter crawling. It avoids IP-based blocking by exploiting public web proxies...

4
WebCrawler2
WebCrawler2 NadavIs56 Python

A simple Python web crawler that processes URLs from web pages, handles redirects, and skips non-HTML content. It supports HTTP/HTTPS, calculates same...

4
ScrapySub
ScrapySub ENGRZULQARNAIN Python

ScrapySub is a Python library designed to recursively scrape website content, including subpages. It fetches the visible text from web pages and store...

4
store-gpt-scraper
store-gpt-scraper apify-projects TypeScript

Extract data from any website and feed it into GPT via the OpenAI API. Use ChatGPT to proofread content, analyze sentiment, summarize reviews, extract...

4
google-image-crawling-extension
google-image-crawling-extension daehwan2 TypeScript

Google Image Auto Download Chrome Extension. 구글 이미지 자동 다운로드 크롬 익스텐션.

4
crawling-from-scratch
crawling-from-scratch ZenRows Python

Repository for the Mastering Web Scraping in Python: Crawling from Scratch blogpost with the final code.

4
Web-Crawling-To-TXT
Web-Crawling-To-TXT fernaerell Python

A simple web crawling application that can browse URLs, extract text content, and save the results in TXT format.

4
SpiderSel
SpiderSel Haxxnet Python

Python 3 script to crawl and spider websites for keywords via selenium

4
STUDY_Python
STUDY_Python Jiyeon1104 Jupyter Notebook

🎈Python 학습 내용을 올린 레파지토리입니다. 🎈

4
scrapy-source
scrapy-source hideaki-kawahara

Sample code for scraping with Python Scrapy.

4