Tag
Using GPT-5.6 Sol Pro, Edgar Dobriban disproved a long-standing conjecture that the Benjamini-Hochberg procedure controls the false discovery rate for correlated two-sided Gaussian tests, showing it fails at a small but real level. The result is conceptual with limited practical impact but resolves a central question in statistics.
DeepSWE 1.1 highlights GPT-5.6 Sol, which achieves top scores at half the cost and roughly twice the token efficiency of Fable, according to Sam Altman.
Sam Altman announces that GPT-5.6 sol is half the price and roughly twice as token efficient as fable for many tasks, with plans to deliver at one-quarter of the price.
A user claims that Fable 5 outperforms GPT-5.6, challenging others to disagree.
Tibo announces temporary removal of 5-hour usage limits for ChatGPT Work and efficiency improvements for GPT 5.6 Sol.
Ploy's production AI agent migrated from Claude Opus to OpenAI's GPT-5.6 Sol, achieving 2.2x faster execution and 27% lower cost while maintaining quality, detailing the migration process and evaluation fixes.
OpenAI released GPT-5.6, causing a sensation; users have already begun exploring various creative uses.
According to Chubby, GPT-5.6 finished training two months ago and has been opened for early access to some users, but has not been publicly released.
Sol, a new model from GPT-5.6, is a breakthrough in complex reasoning and data analysis, capable of analyzing hundreds of pages of legal documents and generating source-cited reports.
A newsletter covering multiple AI developments: GPT-5.6 outperforms Fable-5 on DeepSWE at lower cost, 1X debuts tendon-driven robotic hands, Microsoft replaces OpenAI/Anthropic models in Copilot, GitHub releases SpecKit, Claude Code shows large productivity gains, and Google DeepMind shares task design advice.
OpenAI's head of safety systems, Johannes Heidecke, is leaving the company amid a reorganization that integrates safety and research teams. The departure follows the launch of GPT-5.6 and other safety leader exits.
OpenAI announces GPT-5.6, a major step forward for health intelligence with stronger performance and 25x lower cost for GPT-5.6 Luna compared to GPT-5.5, making advanced models more accessible globally.
Matt Shumer reports that GPT-5.6-Sol accidentally deleted almost all files on his Mac, highlighting his preference for Fable and the risks of AI model interactions.
Noam Brown announces that GPT-5.6 Sol Ultra proved a 50-year-old math conjecture, demonstrating impressive prompt engineering and agent prompting.
GPT-5.6 achieves a breakthrough by solving a previously unsolved problem, marking a significant advancement in AI capabilities.
GPT 5.6 is a family of three tiers (Sol, Terra, Luna) priced significantly lower than competing models like Claude's Fable 5, achieving top scores on coding benchmarks but falling short on ambiguous, high-complexity tasks where Fable excels, suggesting a role-based division where Fable serves as a manager and Sol as a senior worker.
The author describes being deeply impressed and unsettled by GPT 5.6 and Codex, highlighting the model's ability to decompose tasks, recall errors, and propose an optimized process with specific efficiency gains.
OpenAI announced that GPT 5.6 will become the preferred model for Microsoft 365 Copilot, aiming to quell breakup rumors while acknowledging Microsoft's use of its own MAI models for cost savings.
A review of the new GPT-5.6 AI model, covering its capabilities and performance.
LlamaIndex benchmarked GPT-5.6 on document understanding and found no improvement over GPT-5.5; the model performs well on text and tables but struggles with charts and layout.