web crawler / mirror website

Posts 13 of 3 · Page 1 of 1
web crawler / mirror website
Hello,

I wanted to mirror a database to be able to do my research offline too, I tried using HTTrack, which was very slow (two days downloading and one day verifying the files), does anyone know a better way to mirror a database/online encyclopedia? (without crashing their website of course )
you can easily mirror most sites with a scraper... how useful that is questionable
Quote Originally Posted by Dave84311 View Post
you can easily mirror most sites with a scraper... how useful that is questionable
my plan for now would be to a rent cloud server with ubuntu from hetzner and use HTTrack with the setting "Maximum Number of links" on max value (999999999 as example, but be aware that not the wrong websites get downloaded since it scans all links), and then download it from the server to my computer via ftp or I upload it to my Mega account.

- - - Updated - - -

here is my test config from my local vm

I forgot to mention that I wanted to download 100k+ websites, so I want something that can run in the background or on a server since I need to restart my computer due to me testing diffrent drivers on windows 10 and not having a reliable internet connection
Posts 13 of 3 · Page 1 of 1

Post a Reply

Similar Threads

Tags for this Thread

None

Need help?