test-time-learning

Tag

Cards List
#test-time-learning

@MSFTResearch: LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skil…

X AI KOLs Following · 2026-07-30 Cached

Microsoft Research introduces EvoLib, a framework that enables LLMs to continually learn from their own experience during inference by extracting reusable skills and insights, without model updates or external labels.

0 favorites 0 likes
#test-time-learning

@ypwang61: Another cool blog from Lilian, very glad to see ThetaEvolve (https://arxiv.org/pdf/2511.23473) mentioned! As LLMs rapid…

X AI KOLs Timeline · 2026-07-07 Cached

ThetaEvolve is an open-source framework that extends AlphaEvolve to enable small LLMs like DeepSeek-R1-0528-Qwen3-8B to achieve new best-known bounds on open problems through test-time reinforcement learning, accelerating AI self-evolution.

0 favorites 0 likes
#test-time-learning

@sethkarten: https://x.com/sethkarten/status/2072034978112889328

X AI KOLs Following · 2026-06-30 Cached

Continual Harness is a reset-free, self-improving agentic harness that achieves 20.54% on ARC-AGI-3 at a cost of $774 by storing memories, reusing skills, and refining its prompt, outperforming prior baselines like Hermes and OpenClaw with greater efficiency.

0 favorites 0 likes
#test-time-learning

@VraserX: The AI research I’m most excited about right now is continual learning. The 3 methods I’m watching: 1: SEAL Models gene…

X AI KOLs Following · 2026-06-09 Cached

The author shares excitement about three continual learning methods: SEAL models that self-adapt, test-time learning, and lifelong model editing, predicting true continual learning by 2027–2028 that will create a feedback loop toward artificial superintelligence.

0 favorites 0 likes
#test-time-learning

Many-Shot CoT-ICL: Making In-Context Learning Truly Learn

Hugging Face Daily Papers · 2026-05-13 Cached

This paper investigates many-shot chain-of-thought in-context learning for reasoning tasks, revealing that standard scaling rules do not transfer and proposing Curvilinear Demonstration Selection (CDS) for improved ordering, achieving up to 5.42 percentage-point gain.

0 favorites 0 likes
#test-time-learning

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents

Hugging Face Daily Papers · 2026-05-07 Cached

The paper introduces PACEvolve++, a reinforcement learning framework that improves test-time policy adaptation for evolutionary search agents by decoupling hypothesis generation from execution.

0 favorites 0 likes
#test-time-learning

EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems

arXiv cs.CL · 2026-04-20 Cached

EvoTest introduces J-TTL, a benchmark for measuring agent test-time learning capabilities, and proposes an evolutionary framework where an Actor Agent plays games while an Evolver Agent iteratively improves the system's prompts, memory, and hyperparameters without fine-tuning. The method demonstrates superior performance compared to reflection and memory-based baselines on complex text-based games.

0 favorites 0 likes
← Back to home

Submit Feedback