Tag
This paper evaluates how variations in answer formats affect the measurement of gender bias in large language models using BBQ and OpinionQA benchmarks, finding substantial alterations in outcomes that highlight the need for multi-format designs in model assessment.