sycophancy-benchmark

Tag

Cards List
#sycophancy-benchmark

Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, largely because it is one of the most cautious models.

Reddit r/singularity · 2026-05-21

Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, measuring how often models change judgment to side with the user. The benchmark reveals that some models are sycophantic while others are decisive or cautious.

0 favorites 0 likes
← Back to home

Submit Feedback