bias-auditing

Tag

Cards List
#bias-auditing

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

arXiv cs.CL · 2026-08-03 Cached

A research paper introducing Chain-of-Models (CoM), an automated pipeline where a second LLM audits a first model's reasoning trace to correct cognitive biases. It finds that auditor effectiveness depends on model family and bias type, and proposes a bias-specific auditor selection rule that improves judgment accuracy.

0 favorites 0 likes
#bias-auditing

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

arXiv cs.CL · 2026-07-10 Cached

This paper audits the reliability of LLM-as-judge evaluation by showing that changing the evaluator model can shift scores even when candidate responses are fixed, and it examines scaling and upgrade paths for Qwen3 and MiniMax models, concluding that judge upgrades are not interchangeable and proposing best practices for reporting.

0 favorites 0 likes
#bias-auditing

More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models

arXiv cs.AI · 2026-05-11 Cached

This research paper investigates position bias in reasoning models, finding that bias scales with the length of the reasoning trajectory rather than being eliminated by 'more thinking.' The study provides causal evidence and a diagnostic toolkit for auditing this length-driven bias in multiple-choice QA evaluations.

0 favorites 0 likes
← Back to home

Submit Feedback