Hacker Newsnew | past | comments | ask | show | jobs | submit | RataNova's commentslogin

An agent's usefulness is defined by its ability to successfully close a narrow task within the strictly defined boundaries of a sandbox

Expecting a statistic model to follow security rules with ironclad certainty was a pretty naive idea from the start

Yes exactly, it is a crazy engineering decision… if you know what an LLM is.

Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.


The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement

I love how the solution to any software vulnerability in the machine learning world always turns out to be buying even more server hardware from nvidia

If I understand correctly the tool is just trying to impose senior engineering discipline on the agent through a series of pre-commit checks...

The repo lacks hard numbers on token consumption per feature compared to a straightforward prompting session. Without that benchmark measuring whether running ten validation gates actually pays off is impossible

Even if there is a cartel it's secondary now because Nvidia is basically vacuuming up half of Hynix's output for Blackwell and leaving the retail market to fight over the scraps


Half of the new open-source stuff on github is written by claude now, all the way from issues to docs. Models are just vacuuming up this dataset during pretraining, naturally picking up the tone. You don't even need direct distillation via api anymore when the whole internet has turned into one big snapshot of Anthropic's weights


It writes pretty clean code and holds context alright, but it starts stumbling and losing the plot on complex bash scripts with pipelines. Waiting for the weights to drop so we can dig under the hood and see what is going on there


With that logic you shouldnt even deploy to the cloud at all. Someone in AWS can also accidentally share an S3 bucket with backups to the whole internet, there are plenty of precedents for that


> With that logic you shouldnt even deploy to the cloud at all

Given how permissive the architecture of the internet is from TCP/IP on up, I think you could formulate a vigorous argument backing up this assertion.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: