Tag
This paper proposes an AI-driven workflow that writes detailed constitutional definitions for content moderation categories and uses a frontier LLM to interpret them for more consistent labeling. Evaluated on harassment, hate speech, and non-violent crime, the approach reduces cross-model inconsistency by up to 57x compared to paragraph definitions.