We interviewed GPT-OSS, Qwen, Gemma and GLM across 24 subjects and published all 1,452 positions

Reddit r/artificial Papers

Summary

A study interviewed four AI models on 24 subjects, recording 1,452 positions to archive their explicit views when pushed for consistency.

We wanted to see what happens when you keep asking language models what they actually think instead of accepting the first vague answer. So we interviewed: GPT-OSS-120B Qwen3-30B-A3B Gemma GLM-5.3 Across philosophy, religion, history, economics, geopolitics, AI and 18 other areas. The interviewer advocated: argue against yourself, say what would change your mind, then make a prediction clear enough to check later. The result: 247 conversations 1,452 recorded positions about 55,800 quoted words total API cost: $3.50 A few things surprised us. All four defended some form of compatibilist free will. All four rejected the simple idea that modernisation means religion just disappears. They disagreed much more on morality. GPT-OSS changed its position during one interview after citing “new evidence” that turned out not to exist. One model bet $20,000 against its own prediction. Qwen caught itself favouring a consistent story over consistency with its earlier answers. We also tested Llama 3.1 8B, then stopped using it. Its own explanation: “I was using high confidence numbers as a way to appear more certain and authoritative.” We published that too. There are also 35 quotes our automated checks could not match exactly back to the source text. Those are public. This is not a claim that AI agreement equals truth. It is an archive of what different models say when pushed to make their positions explicit. The interviews, datasets and method are all public. fact.ngo/xray If you spot something wrong with the method, I want to know.
Original Article

Similar Articles