Topic

crawler

Repositories (1470)

Craw4LLM
Craw4LLM cxcscmu Python

Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining"

664
DouYin
DouYin Python3WebSpider Python

API of DouYin for Humans used to Crawl Popular Videos and Musics

652
NetDiscovery
NetDiscovery fengzhizi715 Java

NetDiscovery 是一款基于 Vert.x、RxJava 2 等框架实现的通用爬虫框架/中间件。

646
learnPython
learnPython rieuse Python

Python的基础练习代码与各种爬虫代码

646
pywebcopy
pywebcopy rajatomar788 Python

Locally saves webpages to your hard disk with images, css, js & links as is.

639
runoob-PDF-
runoob-PDF- gagayuan Python

爬取菜鸟教程网站并转PDF__python_crawer_by_chrome

638
Moodle-DL
Moodle-DL C0D3D3V Python

Moodle-DL downloads course content fast from Moodle (eg. lecture pdfs)

633
fortress
fortress tiliondev Python

Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change.

631
go_jobs
go_jobs go-crawler Go

带你了解一下Golang的市场行情

620
dotcommon
dotcommon Kharacternyk Python

What do people have in their dotfiles?

620
pryingdeep
pryingdeep iudicium Go

Prying Deep - An OSINT tool to collect intelligence on the dark web.

611
scrapedin
scrapedin linkedtales JavaScript

LinkedIn Scraper (currently working 2020)

610
Jie
Jie yhy0 Go

Jie stands out as a comprehensive security assessment and exploitation tool meticulously crafted for web applications. Its robust suite of features en...

608
newcrawler
newcrawler speed JavaScript

Free Web Scraping Tool with Java

586
mmjpg
mmjpg chenjiandongx Python

👩 美女写真套图爬虫(一)

566
python-automation-scripts
python-automation-scripts avidLearnerInProgress Python

Simple yet powerful automation stuffs.

562
webster
webster zhuyingda JavaScript

a reliable high-level web crawling & scraping framework for Node.js.

560
reader
reader vakra-dev TypeScript

Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.

560
crawljax
crawljax crawljax Java

Crawljax

553
vault
vault abhisharma404 Python

swiss army knife for hackers

551
freeproxy
freeproxy CharlesPikachu Python

FreeProxy: Collecting free proxies from internet. (全球海量高质量免费代理,支持爬取数十个免费代理分享源,支持自定义规则代理筛选,爬虫与数据分析必备,...

547
freshonions-torscraper
freshonions-torscraper dirtyfilthy Python

Fresh Onions is an open source TOR spider / hidden service onion crawler hosted at zlal32teyptf4tvi.onion

540
TorCrawl.py
TorCrawl.py MikeMeliz Python

Crawl and extract (regular or onion) webpages through TOR network

533
Python3Webcrawler
Python3Webcrawler mochazi Python

🌈Python3网络爬虫实战:QQ音乐歌曲、京东商品信息、房天下、破解有道翻译、构建代理池、豆瓣读书、百度图片、破解网易登录、B站模拟扫码登录、小鹅通、荔枝微课

530
nintendo-switch-eshop
nintendo-switch-eshop lmmfranco TypeScript

Crawler for Nintendo Switch eShop

524
opensearchserver
opensearchserver jaeksoft Java

Open-source Enterprise Grade Search Engine Software

517
scrapple
scrapple AlexMathew Python

A framework for creating semi-automatic web content extractors

504
Scan-T
Scan-T nanshihui C

a new crawler based on python with more function including Network fingerprint search

502
Html2Article
Html2Article stanzhai C#

Html网页正文提取

495
scraperai
scraperai scraperai HTML

ScraperAI is an open-source, AI-powered tool designed to simplify web scraping for users of all skill levels.

482
fundus
fundus flairNLP Python

A very simple news crawler with a funny name

474
Fast-Powerful-Whisper-AI-Services-API
Fast-Powerful-Whisper-AI-Services-API Evil0ctal Python

⚡ 一款用于自动语音识别 (ASR)、翻译的高性能异步 API。不需要购买Whisper API,使用本地运行的Whisper模型进行推理,并支持多GPU并发,针对分布式部署进行设计...

471
tsrtc
tsrtc Asoul JavaScript

台灣股票即時爬蟲。Taiwan Stock Exchange Real Time Crawler

462
AllNewsSpider
AllNewsSpider Python3Spiders Python

澎湃新闻,新浪新闻,腾讯新闻,搜狐新闻,新闻联播,泰晤士报,纽约时报,BBCNews,旨在爬取所有新闻门户网站的新闻,禁止将所得数据商用!

462
ICLR2020-OpenReviewData
ICLR2020-OpenReviewData shaohua0116 Jupyter Notebook

Script that crawls meta data from ICLR OpenReview webpage. Tutorials on installing and using Selenium and ChromeDriver on Ubuntu.

460
Antibot-Detector
Antibot-Detector scrapfly JavaScript

Real-time detection of anti-bot systems, CAPTCHAs & fingerprinting techniques. Identifies Cloudflare, Akamai, DataDome, reCAPTCHA, hCaptcha, Shape Se...

452
sitemap-generator
sitemap-generator lgraubner JavaScript

Easily create XML sitemaps for your website.

452
coom-dl
coom-dl notFaad Dart

Coomer| kemono .party or su downloader

438
Pinkerton
Pinkerton dsssssssm Python

🕵️ Python project to crawl for JavaScript files and search for secrets like API keys, authorization tokens, hardcoded credentials, etc.

438
media-scraper
media-scraper elvisyjlin Python

Scrapes all photos and videos in a web page / Instagram / Twitter / Tumblr / Reddit / pixiv / TikTok

433
signature_algorithm
signature_algorithm gadfly0x Python

各种App、小程序、网站的请求签名或加密算法。 现已有:自如、小红书、蛋壳公寓、luckin coffee(瑞幸咖啡)、bangkokair(曼谷航空)

431
Youtube-Projects
Youtube-Projects ayushi7rawat Python

This repository contains all the code I use in my YouTube tutorials.

427
dude
dude roniemartinez Python

dude uncomplicated data extraction: A simple framework for writing web scrapers using Python decorators

426
EmailFinder
EmailFinder Josue87 Python

Search emails from a domain through search engines

425
gospider
gospider nange Go

golang实现的爬虫框架,使用者只需关心页面规则,提供web管理界面。基于colly开发。

415
sosse
sosse biolds Python

Selenium Open Source Search Engine & crawler

411
music-recover
music-recover heqin-zhu Python

:musical_note: 缓存文件转换为 MP3 文件

408
second-order
second-order mhmdiaa Go

Second-order subdomain takeover scanner

408
CrawlerForReader
CrawlerForReader smuyyh Java

Android 本地网络小说爬虫,基于jsoup及xpath

403
magic_google
magic_google howie6879 Python

Google search results crawler, get google search results that you need

403