Tag
This paper presents a framework that augments Large Language Models with geometric vision parsing and symbolic solving to match state-of-the-art multimodal models on complex geometry problems, using a new benchmark from 2025 Chinese Zhongkao exams for evaluation.
This paper introduces Procedural Graph, a framework that organizes LLM agent actions into structured triplets for improved long-horizon tool use, with self-evolving topology to enhance performance.
PragAlign introduces a feedback-guided framework for controlled synthetic dialogue generation using an LLM-based evaluator to improve alignment with intent, emotion, coherence, and fluency, achieving 99.50% acceptance compared to 72.25% for one-shot generation.
EvoUndo introduces a framework for evaluating and ensuring recoverability in self-modifying LLM agents, showing that reliable recovery requires co-designing verification, state grounding, and recovery language expressivity.
The article presents a framework for valuing proprietary data in AI systems, using Google's $10M bid for Spirit's data to calculate the economic uplift needed for break-even.
NVIDIA released Molt, a PyTorch-native agentic RL framework designed for compactness and readability, with performance comparable to Megatron-based stacks. The framework is open-source and includes a paper.
Two Hong Kong students achieved a 5x speedup by adding another loop outside the original automated research framework, without needing a better model or more compute. It is considered one of the most useful papers for ordinary Agent developers.
This project open-sources an AI research system based on the frameworks of value investing masters like Buffett, Munger, Duan Yongping, and Li Lu. It uses Claude Code/Codex to enable multi-agent parallel analysis of financial statements, valuations, etc., and shows real trading returns of over 1.46 million yuan in two years, significantly outperforming major indices.
A detailed guide on building an agentic research framework using a multi-LLM system with persistent memory, allowing researchers to avoid re-explaining context across sessions by leveraging file-based identity, project docs, and memory indices.
This paper introduces PersuasionTrace, a framework for studying multi-turn persuasion in human-LLM interaction, using a Bayesian-network simulated target that models belief updates. The framework reveals that LLMs are persuasive across topics and modalities, and that the Bayesian target better matches human belief dynamics than vanilla LLM simulators.
This week, 9 new records were added to the autoresearch ecosystem, bringing the total to 383, covering multiple open-source tools and projects such as the AutoResearch-RL reinforcement learning framework, lance-autoresearch database kernel optimization, and Clio prediction market backtesting framework.