标签
本文提出了针对自动驾驶汽车中LLM和VLM的新偏见测试方法,表明这些模型在行人礼让决策中继承了基于性别、种族和年龄等属性的人类偏见。
The author presents an off-label evaluation card for Qwen3.6-27B, covering quantization, reasoning mode effects, bias probes, and jailbreak resistance, and compares reasoning effects with Nemotron 3.5 Lightning, finding that thinking mode is net-negative for Qwen but positive for Nemotron.
《华盛顿邮报》对主流AI聊天机器人进行了测试,发现其回答中存在政治偏见,引发对AI系统客观性的担忧。
PNNL 与华盛顿大学的研究人员提出一套系统化框架,测试五种大语言模型在文档中捕捉细微语义变化的能力,揭示位置偏差、上下文连贯效应及模型特有的评分“指纹”。