Tag
This paper introduces a constant-size state cache for block diffusion models, demonstrating significant reductions in memory and latency compared to attention-based approaches, enabling efficient long-context generation without quality loss.
This paper introduces Elo-per-token analysis to study how LLM agents allocate test-time compute, revealing that agents initially outperform independent sampling but slow down over time, with parallel sessions offering performance gains.
SWE-2 is experiencing higher than expected demand, leading to capacity shortages. In response, the free promotion for SWE-2 has been extended into October to accommodate all users.
OpenAI details the evolution of Habitat, their online storage platform, which scaled from a Python library to a service handling over 70 million requests per second to support ChatGPT's billion-user scale.
Neki is a sharded Postgres solution by PlanetScale that enables horizontal scaling to hundreds of millions of QPS and petabytes of data with zero-downtime operations.
SWE-2 is an AI model that matches frontier performance at up to 70% lower cost through scaled reinforcement learning.
Clay relies on LangSmith's online evaluators to scale manual review for millions of agent runs per month and is testing its insights product to better understand agent behavior.
The author argues against a full pause in AI development, emphasizing that responsible advancement is safer than freezing progress, which could shift development to less secure actors, and highlights AI's potential benefits in medicine, clean energy, and more.
Founders from WisprFlow, useactively, and pendoio share how they use Claude Managed Agents' features like outcomes, sandboxing, and memory to scale their products.
A tweet discusses how RLM design principles, including avoiding destructive summarization and enabling model search over context, have remained effective over the past year, referencing a blog post on Codex compaction.
This paper introduces GE-Act 2.0, a world-action model pretrained from scratch to enable scalable zero-shot robotic manipulation with improved success rates across diverse tasks and conditions.
Apodex 1.1 is an AI system that shifts focus from simple answers to completing and verifying entire jobs, with capabilities for environment scaling and agentic coordination.
Jake Castillo shares lessons learned from scaling a consumer app with $50M annual recurring revenue in a new article, sequel to his popular post on building consumer apps.
The article discusses the challenges enterprises face in scaling agentic AI from pilots to full deployment, emphasizing the need for connected strategies, workflow integration, and robust governance to achieve meaningful outcomes.
The paper introduces OR-Transformer, a deep reinforcement learning framework with a permutation-equivariant Transformer architecture for joint replenishment in supply chains, scaling to over 1,000 items and outperforming baselines while reducing decision-making time by millions.
Scaling video pre-training to 120K hours boosts zero-shot success in World Action Models from 36.1% to 77.8% on real robots, enabling faster action prediction.
An interview with EthanrThornton, founder of Mach Industries, discussing starting a solid rocket motor factory, defense industry challenges, manufacturing scaling, and company growth strategies.
This paper examines whether larger batch sizes can reduce wall-clock time-to-target in reinforcement learning for large language models by separating algorithmic and systems-level effects, providing a decision rule for optimization.
OpenClaw 2.0, the fastest-growing open-source project on GitHub, has launched. The article details its evolution from a personal WhatsApp relay to a viral AI agent tool, highlighting the challenges of scaling and maintaining it in the AI era.
This paper proposes a structured ladder for scaling large reasoning models beyond human supervision, addressing challenges in autonomous rewards and self-generated experience, while identifying risks and evaluation dimensions.