human-feedback

Tag

Cards List
#human-feedback

@Vtrivedy10: Synthetic Environment Generation with Human Feedback Agents are poor 1-shot eval/environment generators because they're…

X AI KOLs Following · 2026-08-04 Cached

A tweet introducing the eval-engineering skill from langchain-ai/langchain-skills, which uses human feedback to generate aligned environments, harnesses, and tasks for agent evaluation. It explains the workflow and provides installation instructions for the open-source tool.

0 favorites 0 likes
#human-feedback

DesignArena creators raise $7.9 million to bring taste to AI models

TechCrunch AI · 2026-08-03 Cached

Intelligence, the company behind AI evaluation tool DesignArena, raised $7.9M in seed funding led by Index Ventures. The platform uses human preference rankings to improve AI-generated media and currently has 5.3M users and $60M ARR.

0 favorites 0 likes
#human-feedback

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

arXiv cs.AI · 2026-08-03 Cached

This paper introduces LEMUR, a framework that combines multi-objective reinforcement learning with preference-based learning from multiple human feedback to learn Pareto-optimal policies without predefined reward functions.

0 favorites 0 likes
#human-feedback

Treating "a human rejected this" as a different failure mode than "the agent broke" — turns out that distinction matters a lot in production

Reddit r/AI_Agents · 2026-07-27

A discussion on how treating 'human rejection' as a separate failure mode from 'agent malfunction' significantly impacts the reliability and debugging of AI agents in production.

0 favorites 0 likes
#human-feedback

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

arXiv cs.AI · 2026-07-16 Cached

This paper introduces DROPJ, a human-centred method for safely training and deploying agent policies by learning a world model from real-world trajectories, then eliciting human preferences with justifications to train a reward model for model predictive control. Experiments show that using human-generated simulated trajectories and justifications improves safety and reduces computational cost.

0 favorites 0 likes
#human-feedback

@maxrumpf: Humans Are a Low Ceiling Can Magnus Carlson give feedback to AlphaZero? Obviously not. Even the very best chess players…

X AI KOLs Timeline · 2026-07-07 Cached

Max Rumpf argues that human feedback is becoming obsolete for training advanced AI models, citing examples like chess, math, and search. He advocates for human-free methods like self-play and synthetic data, while a quoted tweet from Will Depue calls for a large-scale data infrastructure parallel to compute scaling.

0 favorites 0 likes
#human-feedback

Building a feedback memory layer for AI agents that learn from every human approval and rejection

Reddit r/AI_Agents · 2026-06-24

This article proposes a feedback memory layer for AI agents that learns from every human approval or rejection, enabling continuous improvement from user interactions.

0 favorites 0 likes
#human-feedback

Hidden Consensus:Preference-Validity Compression in Human Feedback

arXiv cs.CL · 2026-06-10 Cached

This paper argues that standard RLHF's scalarization of human preferences collapses multiple valid interpretations into a single target, mis-measuring alignment in culturally plural societies. Analyzing a Malaysian dataset, they find 79% of prompts have multiple majority-supported responses that single-winner aggregation discards.

0 favorites 0 likes
#human-feedback

What Do People Actually Want From AI? Mapping Preference Plurality

arXiv cs.CL · 2026-06-08 Cached

This paper analyzes 1,500 open-ended responses from 75 countries to reveal that people have diverse and often conflicting preferences for AI, with truthfulness being the only widely demanded value (49%), yet defined in incompatible ways. It argues that current RLHF methods flatten these pluralistic preferences into universal reward models, perpetuating epistemic violence.

0 favorites 0 likes
#human-feedback

StepAudio 2.5 Technical Report

Hugging Face Daily Papers · 2026-05-22 Cached

StepAudio 2.5 is a unified audio-language model that achieves state-of-the-art results across ASR, TTS, and real-time spoken interaction by leveraging task-tailored reinforcement learning from human feedback to optimize shared representations.

0 favorites 0 likes
#human-feedback

@oshaikh13: very cool idea @OpenAI I’m really excited about this research preview- learning from how people interact with their com…

X AI KOLs Following · 2026-04-20

An OpenAI research preview explores learning from how people interact with their computers beyond chat, accompanied by a new arxiv paper on the topic.

0 favorites 0 likes
#human-feedback

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

arXiv cs.CL · 2026-04-20 Cached

WildFeedback is a novel framework that leverages in-situ user feedback from actual LLM conversations to automatically create preference datasets for aligning language models with human preferences, addressing scalability and bias issues in traditional annotation-based alignment methods.

0 favorites 0 likes
#human-feedback

AI-written critiques help humans notice flaws

OpenAI Blog · 2022-06-13 Cached

OpenAI trained language models to write critiques of text summaries, helping human evaluators spot flaws more effectively — a step toward scalable oversight of AI systems on difficult tasks. The work explores how AI-assisted feedback can improve human evaluation quality as a proof of concept for alignment research.

0 favorites 0 likes
#human-feedback

Summarizing books with human feedback

OpenAI Blog · 2021-09-23 Cached

OpenAI presents a scalable alignment technique using hierarchical summarization of entire books with human feedback, demonstrating how models can be trained to act in accordance with human intentions on complex, difficult-to-evaluate tasks.

0 favorites 0 likes
#human-feedback

Learning to summarize with human feedback

OpenAI Blog · 2020-09-04 Cached

OpenAI demonstrates a technique for improving language model summarization by training a reward model on human preferences and fine-tuning models with reinforcement learning, achieving significant quality improvements that generalize across datasets. This work advances model alignment through human feedback at scale, with applications beyond summarization.

0 favorites 0 likes
#human-feedback

Fine-tuning GPT-2 from human preferences

OpenAI Blog · 2019-09-19 Cached

OpenAI demonstrates fine-tuning GPT-2 (774M parameters) using human preference feedback for text continuation and summarization tasks, requiring 5k labels for stylistic tasks and 60k for summarization, with models achieving 86-88% human preference rates though revealing labeler heuristic exploitation.

0 favorites 0 likes
#human-feedback

Learning complex goals with iterated amplification

OpenAI Blog · 2018-10-22 Cached

OpenAI presents iterated amplification, a method for training AI systems on complex tasks by recursively decomposing them into smaller subtasks that humans can judge and solve, building up training signals from scratch through iterative composition.

0 favorites 0 likes
#human-feedback

Gathering human feedback

OpenAI Blog · 2017-08-03 Cached

OpenAI releases RL-Teacher, an open-source tool for training AI systems through human feedback instead of hand-crafted reward functions, with applications to safe AI development and complex reinforcement learning problems.

0 favorites 0 likes
#human-feedback

Learning from human preferences

OpenAI Blog · 2017-06-13 Cached

OpenAI presents a method for training AI agents using human preference feedback, where an agent learns reward functions from human comparisons of behavior trajectories and uses reinforcement learning to optimize for the inferred goals. The approach demonstrates strong sample efficiency, requiring less than 1000 bits of human feedback to train an agent to perform a backflip.

0 favorites 0 likes
← Back to home

Submit Feedback