Tag
Introduces Agentic Monte Carlo (AMC), a method to perform reinforcement learning-style optimization of black-box LLM agents using Sequential Monte Carlo, without requiring access to model parameters.
OpenAI proposes POLO (Plan Online, Learn Offline), a framework combining model-based control with value function learning and coordinated exploration to enable efficient learning on complex control tasks like humanoid locomotion and dexterous manipulation with minimal real-world experience.