emergent-abilities

Tag

Cards List
#emergent-abilities

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

arXiv cs.CL · 2026-08-04 Cached

This paper introduces CurveShift, an analysis method that separates overall ability gains from difficulty-specific improvements in LLM agents. Using METR time-horizon data and LiveCodeBench, it finds that most apparent shifts toward harder tasks are ceiling effects, though a genuine hard-task effect exists for reasoning models in competitive programming.

0 favorites 0 likes
#emergent-abilities

@GoukiMinegishi: Our paper was accepted as a #ICML2026 Spotlight! Reasoning in LLMs has improved largely by chaining local steps. But is…

X AI KOLs Timeline · 2026-05-26 Cached

This paper formalizes analogical reasoning in Transformers using category theory, introduces synthetic tasks to study its emergence, and reveals that it arises from geometric alignment of relational structures and functor application, with signatures also found in pretrained LLMs. The work was accepted as a Spotlight at ICML 2026.

0 favorites 0 likes
#emergent-abilities

@snowboat84: Today, let's discuss something hardcore. One question: what level of mathematics does AI use? From the perspective of tools and models themselves, the mathematics used by AI has an average age of 150 years, with most being from before the mid-19th century: matrix multiplication, gradient descent, chain rule, Fourier transform, inner product, probability — mostly content from the first two years of undergraduate studies. But some phenomena emerging from AI...

X AI KOLs Timeline · 2026-05-23 Cached

Discusses that the mathematics used by AI is mainly linear algebra, calculus, etc., from before the 19th century, but emerging phenomena such as Scaling Law, emergent abilities, double descent, in-context learning, and representation geometry lack mathematical explanation. Analogizes to the clouds in physics in 1900, suggesting it may drive the development of 21st-century mathematics.

0 favorites 0 likes
#emergent-abilities

Your Evals Will Break and You Won't See It Coming

Reddit r/ArtificialInteligence · 2026-05-19 Cached

Discusses the structural weakness of current evaluation methods for LLMs, which fail to anticipate qualitative shifts in capability, and argues that developing proactive evaluation infrastructure is the critical bottleneck for safe capability jumps.

0 favorites 0 likes
← Back to home

Submit Feedback