We interviewed GPT-OSS, Qwen, Gemma and GLM across 24 subjects and published all 1,452 positions
Summary
A study interviewed four AI models on 24 subjects, recording 1,452 positions to archive their explicit views when pushed for consistency.
Similar Articles
@GergelyOrosz: With every new model, I put it to the test in writing - to see if I could do what I do - I give it a bunch of interview…
Gergely Orosz tests new AI models like Fable and GPT-6 by asking them to write articles in his style from his interviews, but finds they fail spectacularly, producing lots of words without understanding.
I tested 6 AI interview assistants and turned the results into a public dataset
A developer tested six AI interview assistant products, found all failed several behavior checks, and published a public dataset with 66 assessments across 6 products and 11 criteria, available on GitHub and Zenodo.
Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest
Researchers from Utah State and Vanderbilt benchmark GPT-4, Gemini 1.5 Pro, DeepSeek-V3, Llama 3.2 and BERT on three social-media tasks—authorship verification, post generation, and user attribute inference—introducing new sampling protocols and taxonomies to reduce bias and enable reproducible benchmarks.
Found a tool that asks GPT, Claude, Gemini, and Grok the same question and gives you one consensus answer
The article highlights AllChat, a tool that queries GPT, Claude, Gemini, and Grok simultaneously and returns a single consensus answer, along with a breakdown of each model's response.
Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]
A solo evaluation of six frontier LLMs on 8 bias benchmarks finds that most models lean left politically, and Grok's self-reported right-leaning stance is inconsistent with its left-leaning behavior. Refusal rates vary, with GPT-5.4 refusing 20% of race-related questions.