But how did you find those sites that had the robot.txt to begin with? LLM must somehow find the existence of those pages and store that information before they can crawl them further or mark as acceptable source.
I think a distinction needs to be made between ingesting for LLM training and ingesting / crawling because a human asked it to during an inference session.
I have been talking about the latter, agree the former is abusive.