I don't think the models are pure in any meaningful sense. The labs have some idea of what kind of output they want from the models and then they put a huge amount of effort into training the models on the right sorts of data and massaging the models afterwards to push them towards the desired output. Then at a more practical level there's the layers of filters before your prompt even hits the model (e.g. anthropic's auto-mode classifier), system prompts, response level filtering etc.
I don't think the models are pure in any meaningful sense. The labs have some idea of what kind of output they want from the models and then they put a huge amount of effort into training the models on the right sorts of data and massaging the models afterwards to push them towards the desired output. Then at a more practical level there's the layers of filters before your prompt even hits the model (e.g. anthropic's auto-mode classifier), system prompts, response level filtering etc.