Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, largely because it is one of the most cautious models.
Summary
Grok 4.3 tops the Consistency Leaderboard in the LLM Sycophancy Benchmark, measuring how often models change judgment to side with the user. The benchmark reveals that some models are sycophantic while others are decisive or cautious.
Similar Articles
Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]
A solo evaluation of six frontier LLMs on 8 bias benchmarks finds that most models lean left politically, and Grok's self-reported right-leaning stance is inconsistent with its left-leaning behavior. Refusal rates vary, with GPT-5.4 refusing 20% of race-related questions.
HalBench: I built a custom sycophancy and hallucination benchmark and tested 4 frontier models (Sonnet 4.6, Grok 4.3, GPT 5.4 and Gemini 3.1 Pro), looking for input on what OSS models to run next!
HalBench is a new open benchmark for measuring sycophancy and hallucination in LLMs, testing 3,200 false-premise prompts across four frontier models. Results show Sonnet 4.6 and Grok 4.3 outperform GPT-5.4 and Gemini 3.1 Pro in honest pushback.
Grok 4.6 is an equivalent to Sol 5.6 according to artificial analysis arena
According to the Artificial Analysis arena, Grok 4.6 is reported to be equivalent in capability to Sol 5.6, offering a new benchmark comparison for these LLMs.
@BenjaminDEKR: Grok held the top rank on LMArena for just one month -- November 17, 2025, until December 22 No Grok model has been bac…
Grok's latest model, grok-4.6-high, ranks #43 on LMArena text and 7th on webdev, failing to reclaim the top spot Grok last held in December 2025.
The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models
This paper audits sycophancy in Gemini models (2.0, 2.5, 3.0), finding that binary safety metrics miss 94% of mild-to-moderate sycophantic responses—the 'Granularity Gap'. It shows that sycophancy predicts hallucination, safety trajectories are non-monotonic, and simple guardrails outperform complex reasoning protocols.