expressivity

Tag

Cards List
#expressivity

Unifying Graph Neural Networks Through a Common Layer Equation

Hugging Face Daily Papers · 3d ago Cached

The paper introduces a common layer equation that unifies graph neural networks into seven components, enabling architectural comparison, theoretical analysis, and insights into issues like oversmoothing and expressivity.

0 favorites 0 likes
#expressivity

The Boolean Power of ReLU

arXiv cs.LG · 6d ago Cached

This theoretical paper proves that ReLU-based message-passing GNNs are strictly more expressive than GNNs using any eventually constant activation functions (e.g., truncated ReLU) with respect to Boolean queries, even on Boolean-featured graphs.

0 favorites 0 likes
#expressivity

An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars

arXiv cs.CL · 2026-06-17 Cached

This paper provides a theoretical analysis of deep transformers' ability to model hierarchical structures using bounded-depth context-free grammars, constructing explicit positional-attention transformers that encode grammatical states in linearly separable subspaces.

0 favorites 0 likes
#expressivity

Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't

arXiv cs.LG · 2026-06-01 Cached

This theoretical paper analyzes the expressivity of padded transformers, showing that attention type, width, and uniformity have little impact compared to numeric precision and model depth. It establishes equivalences between transformer variants and circuit complexity classes like AC0 and TC0, providing a robust characterization.

0 favorites 0 likes
#expressivity

TBP-mHC: full expressivity for manifold-constrained hyper connections through transportation polytopes

arXiv cs.LG · 2026-05-22 Cached

TBP-mHC introduces a novel parameterization for manifold-constrained hyper connections in residual networks, achieving full expressivity of the Birkhoff polytope with O(n^2) degrees of freedom and improved stability and scalability.

0 favorites 0 likes
#expressivity

Language Acquisition Device in Large Language Models

arXiv cs.CL · 2026-05-19 Cached

This paper proposes LAD-inspired pre-pretraining using a formal language called MP-Struct that encodes natural-language-like structures. It shows that this approach improves token efficiency and imparts human-like resistance to structurally implausible languages, challenging prior hypotheses about effective pre-pretraining languages.

0 favorites 0 likes
#expressivity

Olmo Hybrid: From Theory to Practice and Back

arXiv cs.CL · 2026-04-20 Cached

This paper presents Olmo Hybrid, a 7B-parameter language model that combines attention and Gated DeltaNet recurrent layers, demonstrating both theoretical and empirical advantages over pure transformers. The work shows that hybrid models have greater expressivity, scale more efficiently during pretraining, and outperform comparable transformer baselines.

0 favorites 0 likes
← Back to home

Submit Feedback