Classify content safety
check_safetyDetect harmful content such as violence, hate, harassment, self-harm, sexual content, or illegal activity. Get severity ratings, a safety score, and a blocked flag to guide moderation.
Instructions
Classify text for harmful or unsafe content — violence, hate, harassment, self-harm, sexual content, illegal activity, and similar categories. Use this to moderate user-generated content or to screen an AI response before it is shown to a user. Returns the safety categories that were triggered, each with a severity (high / medium / low), plus an overall safe boolean, blocked, and a 0-1 score (1 = clean). High-severity categories mean the content should not be displayed or acted on.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text to classify. |