Tag
SlideBank is a training-free framework for whole-slide image reasoning in pathology, using a persistent hierarchical evidence bank to improve consistency and reduce inference costs.
Mimir v1 is a 1-billion-parameter Hierarchical Reasoning Model trained on ethically sourced data that achieves competitive English performance and sets state-of-the-art results for Danish, matching larger frontier models.
The paper introduces TSR, a framework that decomposes social dialogue into strategic planning and linguistic execution, and LHRL-VGR, a reinforcement learning algorithm with variance-gated rewards. Fine-tuning a Qwen2.5-7B agent with this approach surpasses the GPT-4o baseline by 7.32% in goal completion on the SOTOPIA benchmark.
The paper introduces a knowledge-centric framework for generating ComfyUI workflows by distilling hierarchical knowledge (pseudo-codes, skeletons, strategies) from real workflows and using LLMs to perform reasoning from task descriptions to executable structures, achieving higher node diversity and execution success rates.
Sapient Intelligence has released HRM-Text, a 1B parameter text generation model, trained on only 0.04 trillion tokens (costing approximately $1000), surpassing much larger models trained on 100-1000 times more data on multiple reasoning benchmarks, marking the beginning of a new paradigm for AI training.
HRM-text is a 1B-parameter hierarchical reasoning language model proposed by Sapient Intelligence. It thinks efficiently through internal latent space, achieving performance surpassing most models of the same size with extremely low training cost.
Sapient Intelligence released HRM-Text-1B, a 1-billion-parameter language model with a novel dual-timescale recurrent architecture (Hierarchical Reasoning Model) that provides unbounded compute depth at bounded parameter count. The pre-alignment checkpoint is available on Hugging Face.
The paper proposes ProcedureVQA, a benchmark for visual procedural question answering, and Chain-of-Procedure (CoP), a hierarchical reasoning framework that retrieves relevant instructions using visual cues and refines steps through semantic decomposition, achieving up to 13% improvement over baselines.