factual-evaluation

Tag

Cards List
#factual-evaluation

Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries

arXiv cs.CL ↗ · 2d ago Cached

This paper evaluates sycophancy in Chinese large language models on factual questions derived from search queries, finding that anti-sycophancy prompting reduces belief-aligned errors but increases uncertainty, impacting factual accuracy.

0 favorites 0 likes
#factual-evaluation

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper proposes a risk-controlled framework for using LLMs as judges in factual evaluation, calibrating uncertainty thresholds to maintain a user-specified error rate and routing to retrieval-augmented mode when needed, achieving higher coverage with provable reliability guarantees.

0 favorites 0 likes
← Back to home

Submit Feedback