A simple proxy server to integrate crawl4ai with OpenWebUI
Stateful programmatic web browsing, based on Python-Mechanize, which is based on Andy Lester’s Perl module WWW::Mechanize.
Online search advertisement platform & Realtime Campaign Monitoring [Maybe Deprecated]
app版本今日头条 用户登录/个人主页/关注列表/粉丝列表/评论点赞收藏 关键词搜索 头条视频文章信息/评论数据采集
Crawling the data from lazada, websosanh, compare.vn, cdiscount and cungmua with flexible configs
A script to download images from images.google.com
(2020年最新)斗鱼弹幕抓取及实时弹幕数据可视化,分为crawler(弹幕抓取),server(弹幕统计数据服务器),web(统计数据可视化前端)三部分
:earth_africa: Bittorrent Network Overview through Infohash Indexing, Metadata and IP visualisations of the DHT network
publiccode.yml crawler for the Open Source software catalog of Developers Italia
爬取微博关键词相关信息,有远程开发需求可联系,有需要合作加微信: ytouching
An intelligent web crawler built in Go that extracts email addresses from websites with precision and speed.
📊 Project: Analysis & Data Crawling for Two Football Pages – Manchester United & Liverpool FC ⚽🔍
Integrates Supabase with Crawl4AI and AI Chat to create a powerful web crawling and semantic search solution. Streamlit supabase data visualization. R...
Open-source AI browser agent for web reverse engineering. Turn websites into browser-free API clients and crawlers with CDP, network tracing, and Java...
基于 aiotieba 的高性能百度贴吧异步爬虫工具
Parallel Searching and Crawling Data from PTT 🚀
A tool to generate a static website by crawling the original site.
A web scraper that allows you to fetch Movies and TV series information from the FMovies website. (Download + Stream Links included!)
Supervised text summarization (title generation/recommendation) based on academic paper abstracts, with Seq2Seq LSTM and T5.
A web-crawler to scrape FIFA 20 and 21 players' latest information from Sofifa Website
R integration with Screaming Frog CLI
Simple content crawler for joyreactor.cc
本程序支持关键词搜索、热榜、用户信息、回答、专栏文章、评论等信息的抓取
Comic toolbox, GUI interface, download, reader, translation, conversion, image processing, cropping, compression, convenient for e-book reader users...
🌐 Web scraping script written in python using scrapy library in order to scrape product data from popular Sri Lankan web sites
🔎 Job offers scraper: bulldogjob.pl, indeed.com, it.pracuj.pl, jooble.org, justjoin.it, nofluffjob.com, olx.pl, pracuj.pl, theprotocol.it, useme.com
爬取tumblr关注博主图片
A Moodle Crawler that downloads course content from Moodle (eg. lecture pdfs)
Scraping scripts for various websites.
YT metadata exporter
🤖🍴 A Python script that scrapes UFMG's restaurants menus and publishes them @bot_do_bandejao Twitter profile
The project is a spider that uses scrapy and beautifulsoup4 for crawl picture.
群晖Video Station助手,自动获取豆瓣电影信息,并填写Video Station视频信息
A web spider framework
Network dataset extraction library – part of the KONECT project by Jérôme Kunegis, University of Namur
Spider - web crawler and local wordlist processor to generate frequency sorted wordlist / ngrams
Browser extension that extracts all comments from the YouTube video page, sorts them by the amount of likes and saves them to a csv file.
SpiderMan,a base-on Broadway fast high-level web crawling & scraping framework for Elixir.
An declarative and easy to use web crawler and scraper in C#
A modular web crawling and chat system that allows for ingesting website content through XML sitemaps, converting to vector embeddings, and providing...
This project is a web crawler based on Scrapy, visualization 2D, PageRank
A web crawling tool which tests websites for SSL, Cookies and ADA compliance and also suggests ways to fix them.
Library to crawl and extract internal links from domain
Automated tool for scraping job postings into a .xlsx files inspired by Job Funnel.
:page_with_curl: Scrape football data from Bet365
A Udemy Course Scraper built with bs4 and selenium, that fetches udemy course information. Get udemy course information and convert it to json, csv or...
A Node.js template to be implemented to archive post from any social media.
Serverless Architecture Crawler demo
模拟登录各类网站,操作 API 完成各种不可描述的事情
Pimcore Website Indexer (powered by Zend Search Lucene)