power-law

Tag

Cards List
#power-law

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

arXiv cs.LG · 3d ago Cached

This paper introduces Power Law Graph Attention (PLGA) and the PLDR-LLM architecture, an exact generalization of scaled dot-product attention using input-generated bilinear operators. It presents theoretical results including an inference-collapse theorem, empirical stability measurements, and machine-checked proofs in Lean 4.

0 favorites 0 likes
#power-law

An Interesting Fourier Transform – 1/F Noise

Hacker News Top · 2026-08-07 Cached

This article explains the Fourier transform of power-law functions, focusing on the interesting case of 1/f noise and the symmetry between time and frequency domains.

0 favorites 0 likes
#power-law

A scaling law of contextual persistence in human language

arXiv cs.CL · 2026-07-29 Cached

This paper presents a scaling law showing that the contextual influence of word order in human language decays approximately as 1/d with distance, as measured by the reduction in perplexity from large language models, across multiple languages and corpora.

0 favorites 0 likes
#power-law

Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study

arXiv cs.LG · 2026-07-27 Cached

This paper presents a distributed benchmark study on scaling laws for classical machine learning models on tabular data, showing that power-law fits hold for most model families and quantifying replicator-implementation variance across 127 student runs.

0 favorites 0 likes
#power-law

@rosinality: https://arxiv.org/abs/2606.29858 Why does power-law scaling occur? Loss of individual tokens follows a sigmoidal curve,…

X AI KOLs Timeline · 2026-06-30 Cached

This paper presents a token-level framework showing that power-law scaling in language model loss arises from the aggregation of sigmoidal learning curves of individual tokens, and demonstrates that reshaping training distributions based on token learning times can accelerate validation loss reduction by 11%.

0 favorites 0 likes
#power-law

Scaling Laws, Carefully (25 minute read)

TLDR AI · 2026-06-26 Cached

A comprehensive overview of scaling laws in deep learning, tracing their theoretical roots and empirical findings, and explaining how loss decreases predictably with model size, data, and compute.

0 favorites 0 likes
#power-law

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Hugging Face Daily Papers · 2026-05-28 Cached

This paper investigates the quantitative limits of parametric memory in LLMs using LoRA as a probe, establishing a power law relationship and introducing a threshold-guided optimization method called MemFT for improved memory performance.

0 favorites 0 likes
#power-law

Saturating Scaling Laws for Equational Discovery: A Phenomenology of Growth Dynamics in Three Toy Substrates with Two Real-World Replications

arXiv cs.AI · 2026-05-26 Cached

This paper investigates growth dynamics in deterministic equational discovery across three toy substrates and two real-world replications, finding substrate-conditional saturating power-law scaling.

0 favorites 0 likes
#power-law

Scaling laws for neural language models

OpenAI Blog · 2020-01-23 Cached

Foundational empirical study demonstrating power-law scaling relationships between language model performance and model size, dataset size, and compute budget, with implications for optimal training allocation and sample efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback