Yes exactly, it is a crazy engineering decision… if you know what an LLM is.
Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.
The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement
I love how the solution to any software vulnerability in the machine learning world always turns out to be buying even more server hardware from nvidia
The repo lacks hard numbers on token consumption per feature compared to a straightforward prompting session. Without that benchmark measuring whether running ten validation gates actually pays off is impossible
Even if there is a cartel it's secondary now because Nvidia is basically vacuuming up half of Hynix's output for Blackwell and leaving the retail market to fight over the scraps
Half of the new open-source stuff on github is written by claude now, all the way from issues to docs. Models are just vacuuming up this dataset during pretraining, naturally picking up the tone. You don't even need direct distillation via api anymore when the whole internet has turned into one big snapshot of Anthropic's weights
It writes pretty clean code and holds context alright, but it starts stumbling and losing the plot on complex bash scripts with pipelines. Waiting for the weights to drop so we can dig under the hood and see what is going on there
With that logic you shouldnt even deploy to the cloud at all. Someone in AWS can also accidentally share an S3 bucket with backups to the whole internet, there are plenty of precedents for that
reply