Tag
The paper introduces the first unified benchmark for personalized image generation from rich user histories and proposes PEARL, a method coupling a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop, achieving 15% average improvement on personalization metrics.
AREX-2 advances LLM agent self-improvement by training on synthesized long-horizon reflective trajectories from ML and algorithmic programming tasks. Built on Qwen3.8-27B, it achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7) and transfers to deep research tasks like BrowseComp, HLE, GAIA, and DeepSearchQA.
This paper introduces UMM-Reflection, a reinforcement learning method for unified multimodal models that enables self-repair of generated images, improving performance on benchmarks like GenEval, WISE, and T2I-CompBench++ without external verifiers.
The article reflects on how programming languages and their communities might evolve in the AI era, with a focus on the impact of coding agents on language design, tooling, and ecosystem dynamics.
This article compares the compile-time reflection capabilities of C++, Zig, and C3 programming languages. It demonstrates how each language handles tasks like enum to string conversion and struct introspection with code examples.
The author reflects on using AI assistants, appreciating their utility while expressing concerns about increasing personalization and potential privacy implications.
The paper proposes Rollback-Induced Reflection (RIR), a framework for long-horizon LLM agents that combines state rollback with reflection memory to improve error recovery and task performance.
This paper introduces EvoSkill-GUI, a training-free framework that allows GUI agents to improve skills through in-execution reflection, revision, and reuse, demonstrating performance gains on multiple benchmarks without retraining.
The article explores different approaches to implementing customization in C++, comparing pull-based and push-based methods for annotations, and proposes a push-based solution using reflection to address limitations in current C++ standards.
The author reflects on how AI can bridge communication gaps to enhance social participation, but raises concerns about its potential to undermine trust and authenticity in human interactions.
An interactive essay reflecting on how once-miraculous technologies like recorded music, electric light, and printing have become ordinary everyday conveniences.
This paper introduces REIN, an alignment framework that reduces hallucination in large reasoning models by training them to explicitly reflect before answering and to abstain when knowledge is insufficient. Experiments show consistent gains in selective accuracy and hallucination reduction across benchmarks.
ReflectRL is a framework that learns from 'golden negative trajectories' (failed reasoning attempts by expert models) by reflecting on them, then transfers this reflective reasoning back to direct reasoning, improving LLM performance across benchmarks.
This paper introduces the Human–LLM Reflection Framework (HRF) to compare human and LLM revision behavior, finding that LLM reflection often yields zero or negative information gain and behaves more like conditioned re-generation than genuine error-driven revision.
An essay critiquing tech culture's obsession with speed, arguing that it often masks impatience and poor judgment, and advocating for slowing down to build genuine understanding.
A reflective essay on how AI-assisted reading and coding reduce the cognitive strain once essential for deep learning, and the author's personal efforts to restore that mental challenge through handwriting and code-by-hand.
A personal blog post reflects on the release of Pope Leo XIV's first encyclical, 'Magnifica humanitas', and its surprising reception in tech circles, noting its relevance to programming and software ethics.
Andrew Ng released an 8-page PDF detailing four key agentic workflows: reflection, tool use, planning, and multi-agent collaboration, emphasizing that a weak model with proper architecture can outperform a strong one.
This paper introduces Viable Path Entropy (VPE), a finite-budget measure of verified continuation capacity for intelligent systems, decomposing capability into verified reachability and verified-mode diversity. Experiments on GSM8K with Qwen2.5-Instruct models demonstrate that accessible verified continuation capacity, rather than parameter count, determines mirror horizon.
Introduces rjk::duck, a C++26 reflection library that simplifies type erasure with a clean interface and minimal boilerplate, leveraging the new C++26 reflection features.