Articles from Reddit
The irrationality of ζ(5) has been proved using an AI-assisted method, and the solution is marked as solved on FrontierMath despite not following the Apéry-style proof as initially sought.
Nonobench v1.2 benchmarks 43 LLMs on nonogram puzzles, with GPT-6 Astra achieving the first perfect score on 15x15 puzzles and open-weight models failing the new 20×20 hard mode.
Introducing the first Malleable AI workstation that integrates multiple AI models like Gemini, Claude, Codex, and Grok, allowing users to work across sessions and devices in parallel and build custom widgets with prompts.
A user enthusiastically reviews GPT-6-Luna and Sol, highlighting their superior performance, context management, and cost-effectiveness compared to other AI models and services.
The author discusses how most AI agents increase workload, but highlights the product 'Catch' as it helps reduce tasks by handling reminders, scheduling, and emails without constant supervision.
Claude Opus 5.5 demonstrated autonomous control of a robot arm to replicate Michelangelo's work and independently noticed and corrected a mistake.
The author shares insights from building OSIO, an operating system for human-AI collaboration in businesses, arguing that organizations should persist above individual AI agents to manage capability, authority, and memory.
AI drastically reduces the cost of generating representations like business plans and code, shifting the organizational bottleneck from creation to validation and judgment in real-world contexts.
Scotland may have around 30 AI data centres in the pipeline with a potential electricity demand of 9.7 GW, but many are not publicly registered, raising concerns about transparency and energy planning.
A browser-only open-source tool that connects to an OpenAI-compatible agent and provides a stable, polished UI for testing, with side-by-side performance metrics like spans, turns, and tokens.
Geoffrey Hinton, the 'godfather of AI,' warns that AI could wipe out humanity and suggests tech companies build 'maternal instincts' into AI models to ensure they care about humans.
The author discusses the challenge of verifying the correctness of AI agents' long-term memory and asks for community experiences on tracking and correcting memory errors.
A user shares detailed benchmarking data and personal insights on running AI models locally with varying GPU power limits, evaluating models like gemma4 and qwen3.5 on a modest hardware setup.
The article distinguishes between permission for AI agents to act and the evidence required to justify action, introducing the Organic Intelligence Protocol (OIP) to establish evidence boundaries, illustrated by a recent audit case study.
A Finnish study of over 2,000 workers found no clear link between frequent AI use and workplace exhaustion, but social comparison was consistently associated with higher exhaustion levels.
The article questions why Qwen hasn't released new small-scale models (1B/2B/4B), which affects accessibility and development for hardware-limited users, and mentions a similar issue with Google's Gemma series.
Claude Opus 5.5 has designed a processor that is faster and smaller than the human-made VexRiscv on the HWE benchmark.
herdr is a Rust binary that serves as a background server for running coding agents, praised for its pane state model and agent-driven coordination, but criticized for missing cost accounting and unattended run logging.
Meta launched Muse, a personal AI agent with its own computer, persistent memory, and tools like a terminal and browser. The author tested it by generating a 30-minute video from a single prompt, and it's officially available in the US/Canada with a VPN workaround for India.
US and Russia removed the requirement for human review of AI-generated targets from a UN draft, potentially weakening global efforts to regulate autonomous weapons.