Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is the line between an instruction and a control.

If the agent can reinterpret, edit or relax the rule that constrains it, the rule isn't actually enforcing anything... it's just part of the prompt.

I think the useful split is to tell the agent the rules so it can avoid wasting work, but independently enforce the rules that actually matter.

The agent can decide how to accomplish the task, but it shouldn't also get to decide whether it's authorized to cross the boundary.



The mandatory policy was part of the codebase that the agent was working on. The agent didn’t feel like figuring out how to make the new feature it was working on respect that policy, so it just changed the policy.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: