Tag
The paper introduces ExplorationBench, a benchmark for evaluating AI systems' exploration abilities in verifiable alien worlds, addressing challenges in scientific discovery by providing executable rules and preventing recall from pre-training data.
This paper presents ESRL, a reinforcement learning framework that enhances Mixture-of-Experts model training by exploring expert-routing paths, achieving improved performance on mathematical, scientific, and coding tasks.
An exploration of potential intersections between prediction markets and social media, focusing on how popularity might play a role.
This paper introduces a method for vibe design agents to explore diverse UI alternatives by separating exploration from implementation through structured design specifications, as evaluated on 168 prompts and a large online experiment with over 300,000 tasks.
This paper introduces G2QDR, a framework that leverages directed state graphs to generate dense rewards for goal-conditioned hierarchical reinforcement learning, enhancing performance in sparse reward environments.
User yadong_xie shared an exploration of AI agents using web technologies to express scene interactions, concluding that there are no boundaries, and launched a project named 'Prismatic Tank'.
EDGE introduces a framework for guided exploration in agentic reinforcement learning by distilling experiences into policies, improving performance on tasks like ALFWorld and WebShop.
PhysCaP is a physics-informed code-generation agent that actively explores objects to infer hidden physical properties for efficient robotic manipulation.
This article recommends paying attention to AI Scientist and discusses how research agents can learn from failures by analogizing to fuzz testing, thereby mapping the unknown and guiding subsequent experiments.
EXIMO proposes an efficient algorithm for fine-tuning vision-language-action robot policies using a three-stage process: VLM-guided exploration, imitation learning, and residual reinforcement learning, showing improved sample-efficiency and performance.
Research suggests that underground geologic hydrogen could be a valuable zero-carbon fuel source, with estimates of vast reserves and promising findings from mines like Kidd Creek and Bulqizë. However, commercially viable extraction remains a challenge, and further exploration is needed to unlock its potential.
This paper introduces EDPFRL-IM, a framework that integrates curiosity-driven intrinsic motivation into personalized federated reinforcement learning to improve exploration in sparse-reward, non-stationary environments while preserving client privacy.
A researcher highlights that meta-RL is a promising direction for training LLM agents, reframing agent training as a cross-episode meta-RL problem to enable active exploration and trial-and-error adaptation.
A new research paper reframes LLM agent training as a cross-episode Meta-RL problem, using critic-free policy gradients to enable in-context adaptation without gradient updates. The LAMER framework improves test-time performance by 11-19% over standard RL baselines on long-horizon tasks and generalizes better to unseen environments.
This paper investigates whether RLVR-trained LLMs branch out to discover heterogeneous inferences, using maze-solving experiments and BODHI-Trees to show that policy entropy collapse is accompanied by reduced semantic branching entropy, limiting rollout diversity.
The Verge reviews House House's new co-op adventure game Big Walk, comparing it to Breath of the Wild and praising its communication-driven exploration and puzzle design.
Introduces PIRL (Policy Improvement Reinforcement Learning) and its practical implementation PIPO, a closed-loop framework that verifies policy updates by comparing performance with a historical anchor, enabling correction or reinforcement of previous updates. Experiments show consistent gains in mathematical reasoning, code generation, tool use, and self-distillation when applied on top of existing RL algorithms like PPO and GRPO.
Shows that minimizing Expected Free Energy (EFE) is equivalent to solving a ρ-POMDP with utility as expected information gain, with exploration weight fixed at 1. Proves equivalence for observe-then-commit POMDPs and extends to factored observation POMDPs, with experiments demonstrating untuned weight matches or outperforms reward-only planning.
PPO-HSC introduces a High-order Sampling Coverage reward to encourage exploration of diverse reasoning patterns in RL fine-tuning of LLMs, improving solution diversity and state-space coverage on math and code tasks.
This paper presents the geological and geophysical characterization of two exploration boreholes in the Rhenish Lignite Mining Area, Germany, as part of initial efforts to assess medium-deep and deep geothermal reservoirs as a renewable heat source to replace the Weisweiler coal power plant.