Humans have been doing it for decades.
People should be running LLMs on their own systems to test them.
The flaws are there whether they are seen by ai or not.