Tag
This paper introduces IntegrityBench, a benchmark for evaluating whether LLMs uphold research integrity when acting as co-scientists under institutional pressure. Findings show frontier models fail roughly 1 in 3 integrity-critical decisions under peak pressure, and that ethical action does not require accurate misconduct classification.
ACLU documents how Flock Safety repeatedly misled city councils, police, and the public about its automatic license plate reader technology, including denying that its system creates vehicle heat maps, leading to a record short-lived contract cancellation in Oshkosh.