Glossary
Safety Filter
Safety filter — A layer that inspects prompts or replies and blocks, rewrites or deflects content that falls outside a product's policy.
In more detail
A safety filter is a check applied around the model rather than inside it. It may sit on the incoming message, on the generated reply, or on both, and it can block outright, quietly rewrite, or steer the character into changing the subject.
Filters vary enormously in strictness and in transparency. Some products state their boundaries plainly; others leave users to infer them from refusals, which is a common source of frustration because a silent deflection looks like the character behaving oddly rather than a policy being enforced.
Filters also change without notice. A product that behaved one way in one month can behave differently the next, and this is one of the most frequent complaints in the category.
Some boundaries are not policy choices at all. Content involving minors is prohibited everywhere, and no configuration exists that changes it.