value-function

Tag

Cards List
#value-function

Agentic Monte Carlo: Simulating Reinforcement Learning for Black-Box Agents

arXiv cs.LG · 2026-06-05 Cached

Introduces Agentic Monte Carlo (AMC), a method to perform reinforcement learning-style optimization of black-box LLM agents using Sequential Monte Carlo, without requiring access to model parameters.

0 favorites 0 likes
#value-function

Plan online, learn offline: Efficient learning and exploration via model-based control

OpenAI Blog · 2018-11-05 Cached

OpenAI proposes POLO (Plan Online, Learn Offline), a framework combining model-based control with value function learning and coordinated exploration to enable efficient learning on complex control tasks like humanoid locomotion and dexterous manipulation with minimal real-world experience.

0 favorites 0 likes
← Back to home

Submit Feedback