self-play

Tag

Cards List
#self-play

@maxrumpf: Humans Are a Low Ceiling Can Magnus Carlson give feedback to AlphaZero? Obviously not. Even the very best chess players…

X AI KOLs Timeline ↗ · 2026-07-07 Cached

Max Rumpf argues that human feedback is becoming obsolete for training advanced AI models, citing examples like chess, math, and search. He advocates for human-free methods like self-play and synthetic data, while a quoted tweet from Will Depue calls for a large-scale data infrastructure parallel to compute scaling.

0 favorites 0 likes
#self-play

I made a superhuman Generals.io agent with self-play RL [P]

Reddit r/MachineLearning ↗ · 2026-06-24

Trained a superhuman Generals.io agent using self-play reinforcement learning with a JAX-based pipeline and Vision Transformer. Achieved #1 on human 1v1 leaderboard; all code and a fast JAX simulator open-sourced.

0 favorites 0 likes
#self-play

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

arXiv cs.LG ↗ · 2026-06-24 Cached

EMAgnet introduces parameter-space exponential moving average regularization for policy gradient self-play in large two-player zero-sum games, achieving lower exploitability compared to uniform regularization targets.

0 favorites 0 likes
#self-play

@dair_ai: // Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minute…

X AI KOLs Following ↗ · 2026-06-20 Cached

A research paper that combines a small amount of human demonstrations as a regularization objective with self-play reinforcement learning, enabling human-compatible driving policies using far less human data (30 minutes vs thousands of hours) and training in 15 hours on a single consumer GPU.

0 favorites 0 likes
#self-play

@VukRosic99: A DeepSeek researcher just open-sourced his AutoResearch personal project. For the first time, the AutoResearch Agent a…

X AI KOLs Timeline ↗ · 2026-06-18 Cached

A DeepSeek researcher open-sourced AutoResearch, an autonomous framework that can plan, execute, and debug RL experiments on the DeepSeek 285B model without human intervention, accompanied by a self-play survey paper.

0 favorites 0 likes
#self-play

@teortaxesTex: Deli open sources his AutoResearch.

X AI KOLs Timeline ↗ · 2026-06-17 Cached

Deli Chen open sources his AutoResearch SKILL tool and releases a survey paper on Self-play, inspired by AlphaZero.

0 favorites 0 likes
#self-play

@victor207755822: Deli AutoResearch SKILL is now officially open source! https://victorchen96.github.io/auto_research/framework.html… Alo…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

Deli AutoResearch SKILL is open-sourced, an autonomous framework that automates GPU experiments and RL pipelines, with a companion survey paper on Self-play.

0 favorites 0 likes
#self-play

Discovering Lattice Reduction Strategies via Self-Play

arXiv cs.LG ↗ · 2026-06-16 Cached

This paper presents Delta-Star, a deep reinforcement learning approach using AlphaZero-style self-play to discover superior lattice reduction strategies by interacting with the primitive actions of the LLL algorithm. The learned policy generalizes to higher dimensions and unseen moduli without retraining.

0 favorites 0 likes
#self-play

Self-Play Reinforcement Learning under Imperfect Information in Big 2

arXiv cs.LG ↗ · 2026-05-29 Cached

This paper presents a self-play reinforcement learning framework for the four-player imperfect-information card game Big 2, comparing policy-gradient and value-based methods and finding that PPO with entropy regularization outperforms others.

0 favorites 0 likes
#self-play

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

Hugging Face Daily Papers ↗ · 2026-05-29 Cached

SCOPE is a self-play framework for open-ended tasks that co-evolves a Challenger and Solver policy, achieving up to +10.4 points on benchmarks without external supervision.

0 favorites 0 likes
#self-play

@rohanpaul_ai: Brilliant new paper from Meta, CMU and other labs. Shows that coding agents improve faster by manufacturing their own s…

X AI KOLs Following ↗ · 2026-05-26 Cached

A new paper from Meta, CMU, and other labs presents Self-play SWE-RL, a method where coding agents train themselves by manufacturing and fixing bugs in real codebases, achieving significant gains on SWE-bench benchmarks without relying on human-written tasks.

0 favorites 0 likes
#self-play

CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test

Hugging Face Daily Papers ↗ · 2026-05-22 Cached

CoSPlay is a training-free framework that jointly improves code generation and unit test quality through cooperative self-play, achieving competitive performance without ground-truth unit tests.

0 favorites 0 likes
#self-play

Backprop-free Pong: PC + distributional Hebbian plasticity vs. PPO: 57% vs. 59%, ~1500 lines from scratch [P]

Reddit r/MachineLearning ↗ · 2026-05-19

Explores how close a biologically plausible Hebbian agent can get to PPO on Pong, finding only a 2% gap but identifying catastrophic forgetting under self-play as a key bottleneck.

0 favorites 0 likes
#self-play

A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning

arXiv cs.LG ↗ · 2026-05-19 Cached

This paper identifies a threshold in decision capacity that determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations, showing that eliminating all positive-reach contingent decisions leads to rapid convergence to a deterministic exploitation attractor.

0 favorites 0 likes
#self-play

When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning

arXiv cs.LG ↗ · 2026-05-19 Cached

This paper studies adversarial action masking in self-play reinforcement learning, where an attacker selectively removes legal actions from a victim's action set. The attack is shown to be significantly more damaging than random masking or perturbation baselines across multiple environments and algorithms, and victims do not recover under extended training.

0 favorites 0 likes
#self-play

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play

arXiv cs.AI ↗ · 2026-05-19 Cached

PopuLoRA introduces a population-based asymmetric self-play framework for RLVR post-training of LLMs, where teacher and student LoRA adapters co-evolve to generate increasingly complex problems, overcoming the self-calibration limitation of single-agent self-play.

0 favorites 0 likes
#self-play

@francoisfleuret: Awesome. Seriously, people are harsh with this platform, but if you are careful with whom you follow, it is a constant …

X AI KOLs Timeline ↗ · 2026-05-16 Cached

Eric Jang announces he has been working on a from-scratch implementation of AlphaGo, the 2016 AI breakthrough that inspired him to enter deep learning.

0 favorites 0 likes
#self-play

@neural_avb: This is what you can achieve with 5-6 hours of Self-Play RL training by the way Actors view the projectiles with lidar …

X AI KOLs Timeline ↗ · 2026-05-16 Cached

A thread sharing a video of self-play RL training with lidar and PPO in Unity, followed by a lecture on building AlphaGo from scratch.

0 favorites 0 likes
#self-play

@Michaelzsguo: This is one of the best deep discussions I've seen recently about the fundamentals of reinforcement learning and its relationship to modern AI. Eric Jang and Dwarkesh turned a seemingly retro exercise—rebuilding AlphaGo with today's tools—into a very clear masterclass: why 'search +...'

X AI KOLs Timeline ↗ · 2026-05-15 Cached

A detailed discussion on reinforcement learning and its connection to modern AI, using the reconstruction of AlphaGo with modern tools as a clear example of search and self-play. Key takeaways include neural network amortization of search, credit assignment challenges in LLMs vs AlphaGo, and implications for automated research.

0 favorites 0 likes
#self-play

Self-play helped AI achieve superhuman performance in Go, so why hasn’t it done the same for LLMs? Researchers have found a solution.

Reddit r/singularity ↗ · 2026-05-15

Researchers introduce Self-Guided Self-Play (SGS), a self-play algorithm for LLMs that prevents reward hacking by using a Guide role to score synthetic problems. Applied to theorem proving in Lean4, SGS surpasses RL baselines and allows a 7B model to outperform a 671B model.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback