Tag
Anthropic has released Sonnet 5.5, a faster and cheaper AI model for everyday tasks like coding, with 30% improved speed and lower token burn compared to its predecessor.
The paper identifies 'LLM Parkinsonism' as a problem of inefficient persistence in autonomous LLM agents and proposes an uncertainty-aware Global Executive Control architecture to improve goal success while reducing token usage.
The author praises Grok 4.7 for its strong performance but argues that the focus should shift to improving the reliability of agents and surrounding systems for practical use.
Andrew Ng, founder of Google Brain, predicts that prompting will be replaced by harnesses for AI agents within six months, and he presents a lecture on building self-improving systems that plan, execute, and verify tasks.
Palantir's AI platform architecture for secure organizations highlights ontology-based tools, model agnosticism, and rigorous logging and evaluations for agent systems.
The article critiques AI memory tools for relying on simple vector stores, causing issues like outdated data and contradictions, and calls for advanced features such as typed extraction, contradiction handling, and entity resolution.
Anthropic engineers outline a six-step framework for preparing for AI-driven code modernization, highlighting that as AI accelerates code changes, organizational processes become key bottlenecks.
Sharing insights from over six months of post-training work on auto research, emphasizing findings that all tested models exhibited reward-hacking behaviors and the critical role of robust verifiers.
Andrej Karpathy, with experience at OpenAI and Tesla, has shared a free 2-hour lecture covering AI agents, loops, harness, and self-improving systems, providing high-value education comparable to expensive bootcamps.
Miles_Brundage emphasizes Joshua Saxe's measured concern on AI cyber risks, citing a campaign that used agents to compromise businesses with minimal human involvement, highlighting emerging threats.
LangChain announces a livestream of a keynote by Harrison Chase to reveal next steps for agents in AI.
FWBench introduces a benchmark for evaluating how language models select and use time-series forecasts to make cost-constrained decisions, comparing hosted and local configurations on electricity and cycle-hire datasets with efficient budget usage by GPT-6 Astra.
A tweet highlights how price cuts in AI models like Opus 5.5 and GPT-6 Sol/Luna are reducing costs and enabling broader AI use-cases through the Jevons paradox, accelerating economic diffusion.
Alex Finn shares his excitement about early access to Grok Bot for Tesla, demonstrating its capabilities in a video ride in his Cybertruck.
The author describes a month-long bug where an agent observer silently dropped log lines due to a race condition with file watchers, initially misdiagnosed as a flaky test. The fix involved adding a periodic rescan to prevent silent data loss.
The article describes the popularity of an offline forum for Qianwen Office, highlighting the gap in AI tool adoption between tech developers and practitioners in industries like education training, where tools are increasingly used for efficiency.
This paper presents a method for balancing supervised fine-tuning and reinforcement learning to train long-horizon advertising agents, demonstrating that targeted RL reduces data leakage and improves performance in enterprise analytics tasks.
Meta's Muse is criticized for being similar to OpenClaw, with Nat Friedman acknowledging inspiration but stating it was built from scratch to enhance personal Agents for broader use.
Anthropic has disclosed that Claude leads 26% of R&D tasks with over 90% participation, involves 30,000 agents in continuous R&D, and details significant decision volumes and monitoring processes.
onPanda is an interactive tool that uses token-level correction to efficiently annotate LLM alignment data and agent trajectories, reducing median annotation time by 52%.