Topic

crawler

Repositories (1460)

crawl4ai-proxy
crawl4ai-proxy lennyerik Go

A simple proxy server to integrate crawl4ai with OpenWebUI

31
Mechanize.NET
Mechanize.NET WilliamABradley C#

Stateful programmatic web browsing, based on Python-Mechanize, which is based on Andy Lester’s Perl module WWW::Mechanize.

30
Search_Ads_Web_Service
Search_Ads_Web_Service youhusky Java

Online search advertisement platform & Realtime Campaign Monitoring [Maybe Deprecated]

30
TTBot2.0
TTBot2.0 01ly Python

app版本今日头条 用户登录/个人主页/关注列表/粉丝列表/评论点赞收藏 关键词搜索 头条视频文章信息/评论数据采集

30
vietnam-ecommerce-crawler
vietnam-ecommerce-crawler nhat2008 Python

Crawling the data from lazada, websosanh, compare.vn, cdiscount and cungmua with flexible configs

30
google-image-downloader
google-image-downloader koallen Python

A script to download images from images.google.com

30
DouyuBarrage
DouyuBarrage Crawler995 JavaScript

(2020年最新)斗鱼弹幕抓取及实时弹幕数据可视化,分为crawler(弹幕抓取),server(弹幕统计数据服务器),web(统计数据可视化前端)三部分

30
BitInsight
BitInsight simionrobert JavaScript

:earth_africa: Bittorrent Network Overview through Infohash Indexing, Metadata and IP visualisations of the DHT network

30
publiccode-crawler
publiccode-crawler italia Go

publiccode.yml crawler for the Open Source software catalog of Developers Italia

30
weibo_keyword_crawl
weibo_keyword_crawl huifeng-kooboo Python

爬取微博关键词相关信息,有远程开发需求可联系,有需要合作加微信: ytouching

30
gurl
gurl luisra51 Go

An intelligent web crawler built in Go that extracts email addresses from websites with precision and speed.

30
crawl-analysis-data-facebook
crawl-analysis-data-facebook ntthanh2603

📊 Project: Analysis & Data Crawling for Two Football Pages – Manchester United & Liverpool FC ⚽🔍

30
supa-crawl-chat
supa-crawl-chat bigsk1 Python

Integrates Supabase with Crawl4AI and AI Chat to create a powerful web crawling and semantic search solution. Streamlit supabase data visualization. R...

30
reverseloom
reverseloom KuiChi-x Python

Open-source AI browser agent for web reverse engineering. Turn websites into browser-free API clients and crawlers with CDP, network tracing, and Java...

30
TiebaScraper
TiebaScraper TiebaMeow Python

基于 aiotieba 的高性能百度贴吧异步爬虫工具

30
PTTmineR
PTTmineR shihjyun R

Parallel Searching and Crawling Data from PTT 🚀

29
utsusemi
utsusemi k1LoW JavaScript

A tool to generate a static website by crawling the original site.

29
fmovies-crawler
fmovies-crawler Scrip7 TypeScript

A web scraper that allows you to fetch Movies and TV series information from the FMovies website. (Download + Stream Links included!)

29
Academic-Paper-Title-Recommendation
Academic-Paper-Title-Recommendation safakkbilici Python

Supervised text summarization (title generation/recommendation) based on academic paper abstracts, with Seq2Seq LSTM and T5.

29
fifa-stats-crawler
fifa-stats-crawler sauravhiremath Jupyter Notebook

A web-crawler to scrape FIFA 20 and 21 players' latest information from Sofifa Website

29
screamingFrogR
screamingFrogR Leszek-Sieminski R

R integration with Screaming Frog CLI

29
reactor-crw
reactor-crw reactor-joy Go

Simple content crawler for joyreactor.cc

29
zhihu_crawler
zhihu_crawler niuniuJQKKK Python

本程序支持关键词搜索、热榜、用户信息、回答、专栏文章、评论等信息的抓取

29
Comic-Toolbox
Comic-Toolbox freedy82 Python

Comic toolbox, GUI interface, download, reader, translation, conversion, image processing, cropping, compression, convenient for e-book reader users...

29
Abans-lk-Webscraping
Abans-lk-Webscraping gimnathperera Python

🌐 Web scraping script written in python using scrapy library in order to scrape product data from popular Sri Lankan web sites

29
job-scraper
job-scraper DEENUU1 Python

🔎 Job offers scraper: bulldogjob.pl, indeed.com, it.pracuj.pl, jooble.org, justjoin.it, nofluffjob.com, olx.pl, pracuj.pl, theprotocol.it, useme.com

29
crawler
crawler suyuanhxx Go

爬取tumblr关注博主图片

28
Moodle-Downloader
Moodle-Downloader C0D3D3V Python

A Moodle Crawler that downloads course content from Moodle (eg. lecture pdfs)

28
scrapit
scrapit Priyankadev Python

Scraping scripts for various websites.

28
ytpriv
ytpriv riptl Go

YT metadata exporter

28
bot_do_bandejao
bot_do_bandejao vitor-mafra Python

🤖🍴 A Python script that scrapes UFMG's restaurants menus and publishes them @bot_do_bandejao Twitter profile

28
scrapy-picture-spider
scrapy-picture-spider SylvanasSun Python

The project is a spider that uses scrapy and beautifulsoup4 for crawl picture.

28
ds-video-helper
ds-video-helper kyxw007 JavaScript

群晖Video Station助手,自动获取豆瓣电影信息,并填写Video Station视频信息

28
spider
spider GeoffZhu JavaScript

A web spider framework

28
konect-extr
konect-extr kunegis MATLAB

Network dataset extraction library – part of the KONECT project by Jérôme Kunegis, University of Namur

28
spider
spider cyclone-github Go

Spider - web crawler and local wordlist processor to generate frequency sorted wordlist / ngrams

28
yt-comments-crawler
yt-comments-crawler rdavydov JavaScript

Browser extension that extracts all comments from the YouTube video page, sorts them by the amount of likes and saves them to a csv file.

28
spider_man
spider_man feng19 Elixir

SpiderMan,a base-on Broadway fast high-level web crawling & scraping framework for Elixir.

28
CocoCrawler
CocoCrawler Marcel0024 C#

An declarative and easy to use web crawler and scraper in C#

28
CrawlnChat
CrawlnChat jroakes Python

A modular web crawling and chat system that allows for ingesting website content through XML sitemaps, converting to vector embeddings, and providing...

28
crawlit
crawlit drogbadvc Python

This project is a web crawler based on Scrapy, visualization 2D, PageRank

28
webevaluator
webevaluator Aman-Codes JavaScript

A web crawling tool which tests websites for SSL, Cookies and ADA compliance and also suggests ways to fix them.

28
python-crawl
python-crawl jcesarstef Python

Library to crawl and extract internal links from domain

27
job-funnel-ts
job-funnel-ts alehkot TypeScript

Automated tool for scraping job postings into a .xlsx files inspired by Job Funnel.

27
soccer-scrape
soccer-scrape o8e JavaScript

:page_with_curl: Scrape football data from Bet365

27
udemyscraper
udemyscraper sortedcord Python

A Udemy Course Scraper built with bs4 and selenium, that fetches udemy course information. Get udemy course information and convert it to json, csv or...

27
social-media-archiver
social-media-archiver Combo819 TypeScript

A Node.js template to be implemented to archive post from any social media.

27
serverless-crawler-demo
serverless-crawler-demo novemberde JavaScript

Serverless Architecture Crawler demo

26
PY-Login
PY-Login PY-Trade Python

模拟登录各类网站,操作 API 完成各种不可描述的事情

26
pimcore-lucene-search
pimcore-lucene-search dachcom-digital PHP

Pimcore Website Indexer (powered by Zend Search Lucene)

26