Guardrails are the rules that set what an AI may do on its own: what runs freely, what needs your okay first, and what's denied outright. It's a permission layer, not a tone-of-voice setting.
The failure they exist for isn't malice, it's confident competence pointed at the wrong thing — an agent deciding the cleanest fix for a failing test is to delete the test. When a vendor says their agent has guardrails, the useful question is which bucket each action sits in, and who decides. A rule the model can talk its way past is a suggestion.
