Tag
This paper proposes new bias testing methodologies for LLMs and VLMs in autonomous vehicles, showing that these models inherit human biases in pedestrian yielding decisions based on attributes like gender, ethnicity, and age.
The author presents an off-label evaluation card for Qwen3.6-27B, covering quantization, reasoning mode effects, bias probes, and jailbreak resistance, and compares reasoning effects with Nemotron 3.5 Lightning, finding that thinking mode is net-negative for Qwen but positive for Nemotron.
Testing the GLM 5.2 language model for political bias to assess fairness and neutrality.
The Washington Post tested major AI chatbots and found evidence of political bias in their responses, raising concerns about objectivity in AI systems.
Researchers from PNNL and Washington University introduce a systematic framework to test how five LLMs detect subtle semantic changes in documents, revealing positional bias, context coherence effects, and model-specific scoring fingerprints.