A safety filter is a content-moderation component placed around a generative AI model that screens inputs and outputs to block disallowed, harmful, or policy-violating content. It typically combines classifiers, keyword and pattern rules, and policy thresholds to detect unsafe prompts or generations and refuse, redact, or regenerate them. It is a core safeguard in deployed image and conversational AI products.
Content
- Filters apply classifiers and policy rules at prompt and output stages, refusing or sanitising violations before content reaches users. Tuning thresholds trades off over-blocking against missed harms, and layered filters provide defence in depth around the base model.