Tag
A new paper by Ara Kharazian, tryramp, and RevelioLabs uses firm-level data across 21K U.S. businesses to find that firms adopting AI heavily grow headcount by 10% over two years, contradicting the narrative that AI kills jobs.
This article argues that specialization is inevitable for AI systems, drawing on evidence from optimization theory, evolutionary biology, competitive markets, and machine learning. It interprets a 2026 paper by Goldfeder, Wyder, LeCun, and Shwartz-Ziv to challenge the assumption that greater capability leads to greater generality.
Google deployed an agentic AI peer-reviewer at ICML and STOC conferences, reviewing ~10,000 papers with 30-minute turnaround. The formal paper shows it catches 34% more mathematical errors than zero-shot prompting, setting a precedent for AI-automated scientific review at scale.
This paper introduces Agents-K1, a knowledge graph system built from 2.46 million papers that improves AI agent research by incorporating text, figures, tables, and equations, along with a five-level citation classification. It significantly boosts performance of top models like Gemini-3 and GPT-5.2 on benchmarks, demonstrating that refining knowledge structure can be more effective than scaling model size.
A senior Anthropic engineer published an 11-page paper on Loop Engineering, proposing a new paradigm for building agentic systems centered on feedback loops, isolation, verification, and memory rather than smarter prompts.
Anthropic released an 11-page paper titled 'Loop Design: The Anthropic Playbook for Agentic Systems', arguing that independent verifiers are more critical than prompts in agent design.
A reflection on the landmark 'Attention Is All You Need' paper, highlighting how removing recurrence and relying solely on attention mechanisms revolutionized AI and led to modern LLMs like GPT and Claude.
A 58-page paper from Google DeepMind on building agents specialized in game theory, highlighting key insights from the research.
OpenAI released a new paper "Reinforcement Learning Towards Broadly and Persistently Beneficial Models", proposing the Beneficial Trait RL method, training AI's core traits such as honesty and error correction. After training in the medical domain, performance surged across a wide range of OOD tests, and it can resist malicious fine-tuning, breaking the trade-off between safety and capability.
OpenAI released an open research paper on a method to simulate model deployment using de-identified user requests to anticipate real-world behavior before release.
The author reflects on the paper 'Self-Revising Discovery Systems for Science' which proposes a new agentic architecture using strongly-typed DAGs, schema migrations via Kan extensions, and an MDL gate to distinguish genuine discovery from simple retrieval or search.
Proposes FedSPC, a modular correction method for personalized federated learning that applies control-variate correction only to shared parameters, improving performance across various PFL methods on CIFAR-100 and Tiny-ImageNet.
Natasha Jaques praises the Microsoft MAI-Thinking-1 paper for fully disclosing the training recipe for a frontier model, highlighting the token distribution across pre-training, mid-training, and RL post-training phases, and noting that Yann LeCun's cake analogy was prescient.
This paper introduces Self-Harness, a new paradigm where LLM-based agents iteratively improve their own operating harness—prompts, tools, and control flow—without human engineers or stronger external agents, achieving significant performance gains across multiple models.
This paper introduces a categorical framework for distinguishing genuine scientific discovery from mere retrieval or search in self-improving AI agents, using category theory to formalize regime transitions. The authors demonstrate the framework with a protein mechanics example where an agent's accuracy drops as it tackles harder problems, but its theory compresses more data, indicating real discovery.
PaperMentor is a human-centered multi-agent writing assistant that integrates an expert skill library with specialized agents to provide actionable inline comments on Overleaf, outperforming GPT-5.2 in usability and relevance for AI research papers.
A thread reviewing the paper 'Pretraining Large Language Models with NVFP4' and discussing NVFP4 pre-training, especially for NVIDIA Blackwell.
The author open-sourced Sisyphus Academica, a self-coordinating swarm of over 20 specialized agents for producing publication-ready research papers with novelty engines and adversarial review to avoid hallucinated citations and AI-typical prose.
Introduces Balanced Multimodal Label Reshaping (BMLR), a method that addresses modality imbalance in multimodal learning by reshaping the label space to equalize mapping difficulty across modalities, improving performance across various architectures.
The article presents a published architecture (research paper) for AI companions with persistent state, internal need variables, and memory scoring, seeking investment. The system, PHI // DRIFT, includes 18k+ lines of code and a real-time telemetry dashboard.