Minimax M3 appears to have no political censorship
Summary
The Minimax M3 model appears to have no political censorship, standing out among Chinese LLMs in a bias benchmark.
Similar Articles
Built an political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out. [P]
Researcher built an open-source political compass benchmark with 98 structured questions across 14 policy areas to evaluate frontier LLMs (GPT-5.3, Claude Opus 4.6, KIMI K2). Key finding: refusal patterns and opt-out options significantly shift model positioning, with GPT-5.3 refusing 100% of questions when given an opt-out, while KIMI K2 exhibits topic-specific censorship on Taiwan/Xinjiang despite progressive positions elsewhere.
What political censorship looks like inside an LLM's weights (109 minute read)
This mechanistic interpretability study of Qwen 3.5 uncovers the specific circuit responsible for political censorship, demonstrating how it can be identified, analyzed, and even turned off by steering internal directions. The findings reveal that the model's factual knowledge remains intact, with censorship behavior layered on top.
Defining and evaluating political bias in LLMs
OpenAI presents a comprehensive framework for defining and evaluating political bias in LLMs, introducing a 500-prompt evaluation spanning 100 topics across five bias axes. Results show GPT-5 models achieve 30% bias reduction compared to prior versions, with less than 0.01% of production ChatGPT responses exhibiting political bias.
Minimax M3 vs M2.7
Discussion comparing the new Minimax M3 model to its predecessor M2.7, seeking user feedback after two weeks of release.
after a month with 5 Chinese coding LLMs, is M3 actually going to take the top spot?
A user shares a month-long comparison of five Chinese coding LLMs (Kimi K2.6, GLM-5.1, MiMo V2.5 Pro, MiniMax 2.7, DeepSeek V4 Pro) on a TypeScript/Next.js codebase, rating each in categories like frontend, backend, code review, all-rounder, and reasoning. They note MiniMax 2.7 achieves ~90% of Opus 4.6 quality at ~7% cost and speculate whether the upcoming MiniMax 3.0 will close gaps in planning and test coverage to become the top spot.