Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

the Archive Team are bloody-minded about preserving information

They won't do a good job of that if they encourage too many people to make IP blocks of their machines.



AT, the organization, doesn't own any machines. Operations are all done by members, the majority of whom use consumer ISPs.


Do all of them ignore robots.txt?


Well, when I iterated through the everything2.com namespace and downloaded the majority of their content[1], I respected their robots.txt. After I did that, they changed it to:

  User-agent: * 
  Disallow: /
Does that count?

1: http://bbot.org/blog/archives/2011/01/17/more_fun_with_wget/




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: