WebCrawler2

WebCrawler2

NadavIs56

A simple Python web crawler that processes URLs from web pages, handles redirects, and skips non-HTML content. It supports HTTP/HTTPS, calculates same-domain link ratios, avoids duplicate URLs, and saves results in a TSV file. Designed for easy scalability and future extensions.

4 Stars
0 Forks
1 Watchers
Python Language
31 SrcLog Score
Cost to Build
$628
Market Value
$500

Growth over time

5 data points  ·  2026-04-12 → 2026-08-16
Stars Forks Watchers
💬

How do you feel about this project?

Ask AI about WebCrawler2

Question copied to clipboard

What is the NadavIs56/WebCrawler2 GitHub project? Description: "A simple Python web crawler that processes URLs from web pages, handles redirects, and skips non-HTML content. It supports HTTP/HTTPS, calculates same-domain link ratios, avoids duplicate URLs, and saves results in a TSV file. Designed for easy scalability and future extensions.". Written in Python. Explain what it does, its main use cases, key features, and who would benefit from using it.

Question is copied to clipboard — paste it after the AI opens.

How to clone WebCrawler2

Clone via HTTPS

git clone https://github.com/NadavIs56/WebCrawler2.git

Clone via SSH

[email protected]:NadavIs56/WebCrawler2.git

Download ZIP

Download main.zip

Found an issue?

Report bugs or request features on the WebCrawler2 issue tracker:

Open GitHub Issues