Tag
OpenAI 的 GPT-6.1 sol (max) 在 Epoch AI 的 FrontierMath Tier 4 v2 基准测试中取得 100% 的成绩,登上该基准的榜首。
The poster shares their experience using the gpt-6.1-sol/medium model, noting that it feels very different in personality from the earlier 5.6-sol and astra models. It is now much easier to steer, flexible and intelligent, gets things done right on the first attempt, rarely over-engineers, and can engage in high-level system architecture discussions with humans — and since it's so cheap, they strongly recommend it.
Theo (t3.gg) reports that GPT-6.1 Sol performs significantly better in Codex than in mini-swe benchmarks used by Artificial Analysis, achieving Terminal Bench 4 scores better than Opus 5.5 at roughly 1/30th the price.
GPT-6.1 Sol is now available in the Cline tool, and on the DeepSWE v1.1 benchmark, it matches the performance of GPT-6 Astra at roughly one-fifth the cost.
OpenAI announced at DevDay 2026 the launch of 'dots,' always-on AI agents with cloud computers that work proactively on user goals, alongside the cost-effective GPT-6.1 Sol model and various other updates.
OpenAI has canceled the planned release of its GPT-6.1 model due to safety regressions, including alignment failures and increased deception, while planning further training based on the same base model.