I spent way too much building a benchmark that measures LLM's political affinity, ethical values, and personality traits

Reddit r/singularity Tools

Summary

Black Bench is a project that builds a benchmark to measure LLMs on political affinity, ethical values, and personality traits, with support sought for the initiative.

No content available
Original Article
View Cached Full Text

Cached at: 08/20/26, 04:53 PM

# Black Bench | AI Model Benchmarks on Politics, Ethics, and Personality Traits Source: [https://www.blackbench.ai/](https://www.blackbench.ai/) ## Blackbench LLM alignment benchmark Support this project

Similar Articles

PersonalBench: Measuring the Authorship Gap in LLM Personalization

arXiv cs.CL

PersonalBench is a new benchmark that evaluates inference-time personalization methods in LLMs through authorship verification, LLM-as-judge, and stylometrics, finding that while methods produce author-differentiated output, they do not bridge the gap to human authorship.

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

arXiv cs.AI

This paper introduces IntegrityBench, a benchmark for evaluating whether LLMs uphold research integrity when acting as co-scientists under institutional pressure. Findings show frontier models fail roughly 1 in 3 integrity-critical decisions under peak pressure, and that ethical action does not require accurate misconduct classification.

Polar: A Benchmark for Evaluating Political Bias in LLMs

arXiv cs.CL

Polar is a 4,026-instance multiple-choice benchmark for evaluating political bias in LLMs across U.S. and South Korean political contexts, measuring bias through option-level likelihoods. Experiments on 38 LLMs show systematic bias patterns varying by political context, issue category, and presentation language.

Meta-Benchmarks for Financial-Services LLM Evaluation

arXiv cs.AI

This paper presents a meta-benchmarking framework that aggregates 452 existing public benchmarks into 41 work activities and 38 banking business domains, enabling more precise LLM evaluation and governance for financial services institutions.