Tag
This paper introduces MemProbe, a cognitive-science-inspired framework for evaluating stability-plasticity tradeoffs in agent memory systems through experimental paradigms, providing interpretable profiles of memory maintenance over time.
This paper examines how AI systems can leverage identifiers to infer and reconstruct identities through cross-modal linkage, leading to a collapse in practical obscurity and raising significant privacy concerns.
This paper introduces Reachability-Induced Optimization (RIO) to argue that global optimization claims in AI systems should be based on the actually reachable region, providing theoretical results and benchmark data to support this framework.
Sakeena Fiza, a validation engineer at NVIDIA, ensures hardware systems like the NVIDIA Rubin GPU work correctly at scale before mass production, highlighting the importance of validation in AI infrastructure.
The author argues that recoverability is the real test for autonomous AI agents, highlighting challenges like task persistence and the need for robust recovery mechanisms to ensure true autonomy.
The article discusses a critical automation failure mode where actions succeed but responses time out, leading to duplicates, and advocates for using stable operation IDs and state checks to improve agent evaluation robustness.
The paper introduces LSREP, a longitudinal state-replay protocol for evaluating conversational memory, with ICE v2 as a case study, revealing failure modes through replay and auditing in comparison to vector-RAG.
The author argues that AI agents have a state-integrity problem rather than a memory issue, proposing a State Ledger to distinguish historical facts from current state and track provenance.
AttnFuse introduces a composable DSL for compiling custom attention patterns, including RoPE, to fused GPU kernels, achieving speedups over existing methods like PyTorch's flex_attention.
This paper analyzes how Silicon Valley and Big Tech companies are reshaping the U.S. military-industrial complex through AI-enabled systems and large defense contracts, highlighting issues of transparency and effectiveness.
Φ-Bench is a new benchmark for evaluating large language models on engineering and optimizing AI infrastructure, covering tasks from kernel-level code completion to end-to-end system optimization.
Elon Musk invites users to report concerns about X, while Head of Safety Michael O'Herlihy announces his focus on safety, AI systems, and free expression.
Someone published a complete reference diagram for 10 Claude systems that automate tasks, categorized into information, work, and judgment layers, with details on setup requirements and functionality.
Multi-agent AI systems commonly fail at routing, parallelism, handoffs, and coverage. This post recommends a dispatch matrix, parallel execution, structured handoffs, and a catch-all fallback with logging to fix these issues.
This paper proposes a long-run persistence framework for AI systems using a redundancy-adjusted Artificial Age Score (AAS), showing that indefinite cyclic operation need not lead to unbounded structural aging.
This paper introduces 'evaluation blindness,' a formal framework for silent measurement failures that corrupt AI systems from training to deployment, with case studies, a failure taxonomy validated on 50 real incidents, and a failure budget framework.
This paper proposes a semantic framework to describe AI systems, distinguishing justified claims from misleading outputs, and defines common failures such as hallucination and unsupported assertions.
This is an open-source book called 'Agentic Design Patterns', providing 21 chapters and 7 appendices, systematically explaining AI Agent design patterns, covering from basics to enterprise production environments, with each chapter accompanied by Jupyter Notebook practices.
Ian Channing criticizes companies that waste money on RAG systems that perform no better than Google AI, arguing that access to a corpus doesn't equate to deep expertise and that reasoning cannot be cleanly separated from knowledge.
The Sakana Fugu technical report introduces a trained orchestrator that dynamically selects and coordinates specialist models for tasks, with a faster version (Fugu) and a slower workflow version (Fugu-Ultra) that can design custom teamwork patterns per request.