If Your Guardrail Is a Prompt, You Don't Have a Guardrail
TL;DR
A prompt instruction biases the next-token distribution. It cannot bound it. Real guardrails for agents holding production credentials sit below the prompt, in layers the model cannot read or override: scoped identities, vetted tool surfaces, harness hooks, wire-level statement filtering, provenance-tagged logs, behavioral monito...
评论
?
参与讨论