It uses multiple categories of classification, we use OpenAI's free moderation endpoint https://developers.openai.com/api/docs/guides/moderation they are detailed in that document
@stagas are the criteria for the ai content safety classification mostly generic, like autoblocking vile content publication? haven't (had to) use ai yet so i'm unsure of ai sensibilities in that regard