rlhf

Tag

Cards List
#rlhf

AI alignment is the most important problem we will ever have to face.

Reddit r/artificial ↗ · 14h ago

This post argues that AI alignment is the most critical problem humanity faces, with potential for utopia if solved or catastrophe if not, and critiques current alignment methods as inadequate.

0 favorites 0 likes
#rlhf

@alacheng: TypeSafe AI founder Diogo Almeida, in a speech before the release of Jev, should have realized at OpenAI that aligning …

X AI KOLs Timeline ↗ · 4d ago Cached

TypeSafe AI founder Diogo Almeida discusses the limitations of current AI models and introduces Jev, a new model using RLCD alignment designed for software automation rather than human interaction.

0 favorites 0 likes
#rlhf

@3three_AI: Rather than scrolling through Douyin all night, why not watch the latest speech by Jev founder (former OpenAI researche…

X AI KOLs Timeline ↗ · 4d ago Cached

A speech by the Jev founder, a former OpenAI researcher, presents Jev as the next phase for LLMs, claiming it is 200 times faster, 400 times cheaper, and hallucination-free compared to current methods like RLHF.

0 favorites 0 likes
#rlhf

Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

arXiv cs.CL ↗ · 5d ago Cached

This paper introduces RM-EVAL, a reward model trained on human preference data for reference-free meta-evaluation of grammatical error correction, and shows how it can improve GEC systems via reward-guided text generation.

0 favorites 0 likes
#rlhf

@threeaus: I strongly recommend everyone to watch this speech by Jev founder Diogo Almeida, where he has long predicted everything…

X AI KOLs Timeline ↗ · 6d ago Cached

The article recommends a speech by Jev founder Diogo Almeida, critiquing RLHF in AI models like ChatGPT and Claude Code and advocating for true automation beyond human-preference optimization.

0 favorites 0 likes
#rlhf

Jev by TypeSafe AI reached 13% of teams on Vercel AI Gateway in 24h

Reddit r/ArtificialInteligence ↗ · 6d ago

Jev by TypeSafe AI achieved 13% adoption among teams on Vercel AI Gateway within 24 hours, underscoring the importance of seamless integration into existing workflows for rapid product uptake.

0 favorites 0 likes
#rlhf

A new kind of AI model from a ChatGPT inventor is thrilling developers

TechCrunch AI ↗ · 2026-09-18 Cached

A new AI model named Jev, developed by a ChatGPT inventor from TypeSafe AI, outputs probabilities instead of text, offering a fast, cheap, and hallucination-free alternative for software automation tasks.

0 favorites 0 likes
#rlhf

Learning Heterogeneous Preferences

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper introduces a multi-stage architecture for learning individuated utility functions from multi-modal data, showing that models accounting for heterogeneous preferences outperform universal utility models in subjective tasks like aesthetic judgments.

0 favorites 0 likes
#rlhf

@shiri_shh: babe wake up The guy who co-created ChatGPT and RLHF at OpenAI just launched a new model - Latency: ~150ms - Cost: $0.0…

X AI KOLs Following ↗ · 2026-09-15 Cached

Diogo Almeida, co-inventor of ChatGPT, launches Jev, a new frontier AI model trained with RLCD, featuring ~150ms latency and free output cost.

0 favorites 0 likes
#rlhf

@suraj_sharma14: If people in your network need money while they job-hunt, this is the 2026 map. AI labs pay humans to train and grade m…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

The article provides a guide to finding paid remote AI work, listing platforms where humans are hired to train and grade AI models, and advises applying to multiple opportunities.

0 favorites 0 likes
#rlhf

@ProfTomYeh: RLHF by hand ~ 15 steps walkthrough below Train a model on human text and it inherits human bias. It will assume a doct…

X AI KOLs Timeline ↗ · 2026-08-28 Cached

A step-by-step walkthrough explaining how Reinforcement Learning from Human Feedback (RLHF) corrects bias in AI models, using a simple example where a single human preference about doctors generalizes to other professions like CEOs.

0 favorites 0 likes
#rlhf

Why I believe AGI requires inverting current architecture: The Dreamer & The Scribe

Reddit r/ArtificialInteligence ↗ · 2026-08-23

An independent observer argues that achieving AGI requires inverting current AI architecture by combining a stochastic 'Dreamer' with a deterministic 'Scribe', critiquing token-based models, passive learning, and RLHF.

0 favorites 0 likes
#rlhf

@ProfTomYeh: Kimi 3 seminar recording is uploaded http://byhand.ai/v/kimi3 ~ Prof. Tom Yeh

X AI KOLs Timeline ↗ · 2026-08-21 Cached

Prof. Tom Yeh uploads a seminar recording on Kimi 3, featuring special guest Nathan Lambert, with discussions on RLHF, model architecture, and frontier AI topics.

0 favorites 0 likes
#rlhf

@FinanceYF5: Tonight, skip a TV show and finish this 2-hour 34-minute Stanford course. It covers from Tokenization, BPE to Transformer, pre-training, RLHF, DPO, and token-by-token generation, fully deconstructing how large models like ChatGPT and Claude are built…

X AI KOLs Following ↗ · 2026-08-20 Cached

A recommended Stanford course on AI that details the principles behind building large language models, covering Tokenization, BPE, Transformer, pre-training, RLHF, and DPO.

0 favorites 0 likes
#rlhf

Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper defines and quantifies sentiment drift in RLHF-trained summarization models, proposes a Policy Attribution framework to identify causes, and introduces a Sentiment-Aware KL Regularization method to reduce drift.

0 favorites 0 likes
#rlhf

@seclink: Keywords, interested friends can query themselves: RLHF -> RLAIF -> RLTHF. Starting in 2025, RLTHF is adopted, reducing expert working hours by 93%.

X AI KOLs Timeline ↗ · 2026-08-15 Cached

The article briefly introduces the evolution of AI training methods from RLHF to RLAIF to RLTHF and predicts that by 2025, RLTHF will significantly reduce expert working hours.

0 favorites 0 likes
#rlhf

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv cs.LG ↗ · 2026-08-12 Cached

This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.

0 favorites 0 likes
#rlhf

A question about Large Language Models (LLMs): my own observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-08

A user reports that a long, non-instructional text prefix can shift LLM activations and bypass RLHF safety constraints without adversarial prompting, asking whether this reflects distinct world regions in the model.

0 favorites 0 likes
#rlhf

Independent LLM "research" & a direct message to Anthropic ; Preliminary observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-06

An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.

0 favorites 0 likes
#rlhf

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

arXiv cs.LG ↗ · 2026-08-05 Cached

SMOPD proposes a two-stage specialize-and-merge online policy distillation method to improve multi-reward reinforcement learning, addressing issues with sparse and dense reward signals where GDPO struggles. It outperforms GDPO across 1.5B, 3B, and 7B backbones in complementary and conflicting reward settings.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback