The Scaling Properties of Implicit Deductive Reasoning in Transformers
Summary
This research examines how deep Transformers with bidirectional masking achieve implicit deductive reasoning comparable to explicit chain-of-thought methods. The study demonstrates that algorithmically aligned models can scale reasoning capabilities across diverse graph topologies and problem widths.
View Cached Full Text
Cached at: 05/08/26, 02:27 PM
Paper page - The Scaling Properties of Implicit Deductive Reasoning in Transformers
Source: https://huggingface.co/papers/2605.04330 Published on May 5
·
Submitted byhttps://huggingface.co/envomp
Enricoon May 8
Abstract
Deep Transformers with bidirectional masking exhibit implicit deductive reasoning capabilities comparable to explicit chain-of-thought methods across various graph structures and problem sizes.
We investigate the scaling properties ofimplicit deductive reasoningoverHorn clausesindepth-bounded Transformers. By systematically decorrelating provability from spurious features and enforcingalgorithmic alignment, we find that in sufficiently deep models with abidirectional prefix mask, implicit reasoning approaches explicit CoT performance across graph topologies and problem widths, though CoT remains necessary for depth extrapolation.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2605\.04330
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.04330 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.04330 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.04330 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
This paper introduces TranSGrid, a testbed that integrates deductive, inductive, and abductive reasoning to evaluate systematic generalization in AI. Experiments with Transformers show that current tasks overlook essential reasoning aspects, resulting in performance gaps on the proposed testbed.
@yingfan_bot: New paper on Looped Transformers! Latent reasoning is fast, but struggles to match CoT-level accuracy at scale. Can loo…
A new paper on Looped Transformers finds that a looped padded backbone provides a parallel workspace for latent reasoning, enabling supervision similar to explicit chain-of-thought (CoT) and achieving both speed and accuracy.
@machinestein: ICML 2026: Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Why does recursive reasoning, especially …
The paper reveals that latent reasoning in transformer-based reasoning models (TRMs) functions as a policy improvement operator, and proposes an algorithm that enhances learning and inference efficiency by up to 18x.
Syntax vs. Semantics: How Transformers Learn Deep Dependencies
This paper introduces a mechanistic framework analyzing transformer learning dynamics, identifying gradient starvation as a barrier to deep semantic dependencies and validating chain-of-thought strategies for effective learning.
Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models
This paper evaluates the capability and efficiency of large language models using Chain-of-Thought reasoning, finding that capability gains diminish with increasing model size, while efficiency shows little improvement.