Topic

crawler

Repositories (1456)

scrapy_helper
scrapy_helper facert CSS

Dynamic configurable crawl (动态可配置化爬虫)

87
extension
extension get-set-fetch TypeScript

web scraping extension

87
twitter_user_tweet_crawler
twitter_user_tweet_crawler kaixinol Python

A Python crawler tool that can automatically simulate browser operations to crawl all users' tweet content and save all static resources (videos, pict...

87
ICLR2023-OpenReviewData
ICLR2023-OpenReviewData fedebotu Jupyter Notebook

Crawl & Visualize ICLR 2023 Data from OpenReview

87
webspot
webspot crawlab-team Python

An intelligent web service to automatically detect web content and extract information from it.

86
Bilibili_manga_download
Bilibili_manga_download Randark-JMT Python

带图形界面的哔哩哔哩漫画下载工具

86
bilibili-manga-download-script
bilibili-manga-download-script lanyeeee TypeScript

一个用于 哔哩哔哩漫画 B漫 的下载脚本

85
shopify-app-store-scraper
shopify-app-store-scraper usernam3 Python

Crawler behind the Shopify App Marketplace dataset

85
is-google
is-google roccomuso JavaScript

Verify that a request is from Google crawlers using Google's DNS verification steps

84
GMaps-Crawler
GMaps-Crawler guilatrova Python

Google Maps crawler using Selenium. All extracted data is forwarded to a SQS queue.

84
skweez
skweez edermi Go

Fast website scraper and wordlist generator

84
ipgw-py-manager
ipgw-py-manager Neboer Python

NEU new ipgw python manager

83
xSMTP
xSMTP aziz0x48 Python

xSMTP 🦟 Lightning fast, multithreaded smtp scanner targeting open-relay and unsecured servers in multiple network ranges.

83
Hands-on-WebScraping
Hands-on-WebScraping superryeti Python

This repo is a part of blog series on several web scraping projects where we will explore scraping techniques to crawl data from simple websites to we...

83
metacritic_api
metacritic_api melroy89 PHP

PHP Metacritic API - Mirror from my GitLab

82
crawlzone
crawlzone crawlzone PHP

Crawlzone is a fast asynchronous internet crawling framework for PHP.

82
page-redirect-code-location-hook
page-redirect-code-location-hook JSREI HTML

JS逆向技巧:页面跳转JS代码定位通杀方案

81
instastories-backup
instastories-backup ondrejsojka Python

Backup your friends' Instagram Stories forever and get to keep them even after 24 hours.

81
car-prices
car-prices go-crawler Go

Golang爬虫 爬取汽车之家 二手车产品库

81
YourLesson
YourLesson Lewin671 Python

深圳大学抢课系统

80
achoz
achoz kcubeterm Python

Search through all your personal data efficiently like web search.

80
ClaudeChrome
ClaudeChrome NatsuFox JavaScript

ClaudeChrome - Native browser context awareness for agents.

80
simpyder
simpyder Jannchie Python

超高速异步协程Python爬虫

80
open-gov-crawlers
open-gov-crawlers public-law Python

Parse government documents into well formed JSON

80
tieba-zhuaqu
tieba-zhuaqu ankanch Python

百度贴吧分布式爬虫,用于贴吧数据挖掘。从贴吧维度和用户维度进行数据分析

80
arachnid
arachnid watzon Crystal

Powerful web scraping framework for Crystal

79
fund-crawler
fund-crawler nullpointer TypeScript

基于NodeJS的基金数据爬虫,爬取的数据存于github的@nullpointer/fund-data。

79
xueqiu_spider_LQH_LZQ
xueqiu_spider_LQH_LZQ 1491270550 Python

雪球爬虫 高效爬取近期沪深A股股票评论并自动生成PDF版情感分析报告

79
puppeteer-walker
puppeteer-walker lrlna JavaScript

a puppeteer walker 🕷 🕸

79
ctrip_spider
ctrip_spider evanleungc Python

Scrape Learning (ctrip)

79
scrapy-examples
scrapy-examples feiskyer Python

Some scrapy and web.py exmaples

79
slideshare-downloader
slideshare-downloader yodiaditya Python

Slideshare to PDF downloader. Using Selenium and auto scroll-down to get the entire slides completely.

79
ForgeRSS
ForgeRSS tmwgsicp Python

将任意网站转换为 RSS 订阅源、多引擎抓取、反爬突破、支持爬取抖音、快手、小红书、B站、知乎、小宇宙、知识星球

78
ceiba-dl
ceiba-dl lantw44 Python

NTU CEIBA 資料下載工具

78
fetchman
fetchman da2vin Python

fetchman is a simple crawler system/简单好用的爬虫框架

78
tumblr_crawler
tumblr_crawler 2024baibai Python

tumblr解析网站

78
Google-Patents-Scraper
Google-Patents-Scraper wenyalintw Python

Automatically download all PDF files of searching results & their patent families found on Google Patents.

77
cache-warmup
cache-warmup eliashaeussler PHP

🔥 PHP library to warm up caches of URLs located in XML sitemaps

77
tw-stock-telegram-bot
tw-stock-telegram-bot x3388638 JavaScript

台股機器人,提供即時個股及大盤報價、走勢、新聞、盤後資料等 Telegram bot to query real-time TW stock quotes, charts, news, and other related informatio...

77
Pasta
Pasta Kr0ff Python

A PasteBin scrapper that doesnt rely on the PasteBin scrape API

77
tg_crawler
tg_crawler vhdmsm Python

Just a crawler based on tg-cli for Telegram. Deprecated by now, please use telegram-export.

77
eastmoney
eastmoney minicloudsky JavaScript

python requests + Django+ nodejs koa+ mysql to crawl eastmoney fund and stock data,for data analysis and visualiaztion .

76
social-scraper
social-scraper behitek Python

Vietnamese text data crawler scripts for various sites (including Youtube, Facebook, 4rum, news, ...)

76
venom
venom PreferredAI Java

Your preferred open source focused crawler for the deep web.

76
steam-discount
steam-discount EXP-Tools Python

steam 特惠游戏榜单(自动刷新)

76
awesome-fingerprinting
awesome-fingerprinting embeddinglayer

A collection of browser fingerprinting projects, research, and resources. Intended as a way to aggregate research surrounding the subject.

76
midnight_sea
midnight_sea RicYaben Python

Midnight Sea: navigating in the waters of dark web markets

76
librengine
librengine liameno C++

Privacy Web Search Engine (not meta, own crawler)

75
Website-Crawler
Website-Crawler pc8544 Java

Extract data from websites in LLM ready JSON or CSV format. Crawl or Scrape entire website with Website Crawler

75
JobApplicationBot
JobApplicationBot drkostas HTML

A bot that automatically sends emails to new ads posted in any desired xe.gr search url.

75