Tag
The name 'gpt-6-astra-aeon' is confirmed for a new long-running persistent AI agent, indicating a significant update in AI model development.
Reflects on the continuity problem for long-running AI agents, arguing that a deterministic control layer is needed to manage authoritative state, and questions whether existing infrastructure like IAM, transactions, and provenance is sufficient.
LoopX is an open-source control plane for ultra-long-horizon AI agents. By externalizing structured state (todo, authority, evidence, gate, etc.), it enables agents to run continuously for 200+ hours without memory loss or drift, and uses an executable Kanban and a six-layer architecture to manage long-horizon tasks.
An exploration of strategies and techniques used to manage and reduce costs for long-running AI agent deployments.
Anthropic published a detailed guide on how to make Claude work autonomously for 6 hours, covering failure modes, testing with Puppeteer MCP, and session protocols.
A cloud AI agent reportedly ran continuously for 107 hours, highlighting advances in persistent agent operation.
A user let a Fable 5 agent run continuously for 6 days without human intervention, concluding that most people only use 10% of its capacity.
Google shares a free, comprehensive example of a long-running AI agent that pauses, resumes, and never loses context, simulating new employee onboarding, teaching three architectural patterns.
A thread sharing practical tips for running AI agents autonomously for extended periods, focusing on the Opus model with advice on permissions, dynamic workflows, and verification.
Practical tips for running Anthropic's Claude Opus autonomously for hours or days, such as using auto mode, dynamic workflows, and self-verification; also references the SWE-Marathon benchmark for long-horizon software tasks.
NVIDIA introduces Nemotron 3 Ultra, a new AI model designed to enable faster and more efficient reasoning for long-running agents.
Mistral Vibe is an AI agent designed for long-running, multi-step work and coding tasks.
Ivan Burazin notes that 2.5% of sandboxes running over 24 hours generate 20% of revenue, arguing that long-running stateful workloads are common, not edge cases.