Tag
The paper introduces RA-DPO, a reliability-aware direct preference optimization method that combines annotator agreement, model confidence, and token-level uncertainty for sexism detection, improving training efficiency and enabling selective prediction.
The paper proposes MAP-PO, a multi-agent framework that clusters annotators by labeling behavior and fine-tunes separate LLM agents per cluster using preference optimization, preserving disagreement in sexism detection tasks. Experiments on the EXIST 2024 dataset show that cluster-specific training is necessary and that a shared team-level reward keeps agents calibrated.