@hud_evals: Announcing HUD's RL environments for RSI hackathon! Join us June 20–21 in SF if you're interested in RL and want to pus…
Summary
Announcement of HUD's RL environments for the RSI hackathon in San Francisco on June 20-21, offering over $100,000 in prizes and compute credits.
View Cached Full Text
Cached at: 06/01/26, 01:03 AM
Announcing HUD’s RL environments for RSI hackathon! 🎉
Join us June 20–21 in SF if you’re interested in RL and want to push the frontier forward! (w/$100,000+ in prizes and compute credits 👀) https://t.co/GRZfvUGXfe
Similar Articles
Agent 0 from AI-2027 is here - it's called Astra.
An article speculating about OpenAI's internal frontier model Astra (Agent 0), trained with minimal compute, and predicting that future models like Agent-1 will cause widespread job disruption by early 2027.
@rohanpaul_ai: New Meta Paper. Code optimization looks like an easy extension of reinforcement learning: reward correct programs, then…
A Meta paper analyzes why standard RL recipes fail for code optimization and rebuilds the entire feedback pipeline with calibrated timing, problem-relative ranking, and GRPO changes, improving Qwen 2.5 7B speed threshold from 18.0% to 31.3%.
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning
This paper introduces KGPS, a Kalman-guided prompt selection method for adaptive RL finetuning of LLMs, which models prompt difficulty as a dynamic state to improve accuracy and rollout efficiency.
Policy Gradient Steering: Interventions from Behavioral Objectives
Introduces Policy Gradient Steering (PGS), a method that formulates activation steering as a reinforcement learning problem, using policy gradients to construct removable, composable steering vectors from behavioral objectives. Validated in gridworld, chess puzzle, and football environments.
FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning
FinSMART introduces a market-aligned reinforcement learning framework for financial sentiment analysis, optimizing sentiment signals with realized market outcomes and achieving a 220% improvement in cumulative trading returns over the strongest baseline.