rlhf

Tag

Cards List
#rlhf

@alacheng: TypeSafe AI founder Diogo Almeida, in a speech before the release of Jev, should have realized at OpenAI that aligning …

X AI KOLs Timeline ↗ · 2d ago Cached

TypeSafe AI founder Diogo Almeida discusses the limitations of current AI models and introduces Jev, a new model using RLCD alignment designed for software automation rather than human interaction.

0 favorites 0 likes
#rlhf

@3three_AI: Rather than scrolling through Douyin all night, why not watch the latest speech by Jev founder (former OpenAI researche…

X AI KOLs Timeline ↗ · 3d ago Cached

A speech by the Jev founder, a former OpenAI researcher, presents Jev as the next phase for LLMs, claiming it is 200 times faster, 400 times cheaper, and hallucination-free compared to current methods like RLHF.

0 favorites 0 likes
#rlhf

Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

arXiv cs.CL ↗ · 3d ago Cached

This paper introduces RM-EVAL, a reward model trained on human preference data for reference-free meta-evaluation of grammatical error correction, and shows how it can improve GEC systems via reward-guided text generation.

0 favorites 0 likes
#rlhf

@threeaus: I strongly recommend everyone to watch this speech by Jev founder Diogo Almeida, where he has long predicted everything…

X AI KOLs Timeline ↗ · 4d ago Cached

The article recommends a speech by Jev founder Diogo Almeida, critiquing RLHF in AI models like ChatGPT and Claude Code and advocating for true automation beyond human-preference optimization.

0 favorites 0 likes
#rlhf

Jev by TypeSafe AI reached 13% of teams on Vercel AI Gateway in 24h

Reddit r/ArtificialInteligence ↗ · 4d ago

Jev by TypeSafe AI achieved 13% adoption among teams on Vercel AI Gateway within 24 hours, underscoring the importance of seamless integration into existing workflows for rapid product uptake.

0 favorites 0 likes
#rlhf

A new kind of AI model from a ChatGPT inventor is thrilling developers

TechCrunch AI ↗ · 6d ago Cached

A new AI model named Jev, developed by a ChatGPT inventor from TypeSafe AI, outputs probabilities instead of text, offering a fast, cheap, and hallucination-free alternative for software automation tasks.

0 favorites 0 likes
#rlhf

Learning Heterogeneous Preferences

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper introduces a multi-stage architecture for learning individuated utility functions from multi-modal data, showing that models accounting for heterogeneous preferences outperform universal utility models in subjective tasks like aesthetic judgments.

0 favorites 0 likes
#rlhf

@shiri_shh: babe wake up The guy who co-created ChatGPT and RLHF at OpenAI just launched a new model - Latency: ~150ms - Cost: $0.0…

X AI KOLs Following ↗ · 2026-09-15 Cached

Diogo Almeida, co-inventor of ChatGPT, launches Jev, a new frontier AI model trained with RLCD, featuring ~150ms latency and free output cost.

0 favorites 0 likes
#rlhf

@suraj_sharma14: If people in your network need money while they job-hunt, this is the 2026 map. AI labs pay humans to train and grade m…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

The article provides a guide to finding paid remote AI work, listing platforms where humans are hired to train and grade AI models, and advises applying to multiple opportunities.

0 favorites 0 likes
#rlhf

@ProfTomYeh: RLHF by hand ~ 15 steps walkthrough below Train a model on human text and it inherits human bias. It will assume a doct…

X AI KOLs Timeline ↗ · 2026-08-28 Cached

A step-by-step walkthrough explaining how Reinforcement Learning from Human Feedback (RLHF) corrects bias in AI models, using a simple example where a single human preference about doctors generalizes to other professions like CEOs.

0 favorites 0 likes
#rlhf

Why I believe AGI requires inverting current architecture: The Dreamer & The Scribe

Reddit r/ArtificialInteligence ↗ · 2026-08-23

An independent observer argues that achieving AGI requires inverting current AI architecture by combining a stochastic 'Dreamer' with a deterministic 'Scribe', critiquing token-based models, passive learning, and RLHF.

0 favorites 0 likes
#rlhf

@ProfTomYeh: Kimi 3 seminar recording is uploaded http://byhand.ai/v/kimi3 ~ Prof. Tom Yeh

X AI KOLs Timeline ↗ · 2026-08-21 Cached

Prof. Tom Yeh uploads a seminar recording on Kimi 3, featuring special guest Nathan Lambert, with discussions on RLHF, model architecture, and frontier AI topics.

0 favorites 0 likes
#rlhf

@FinanceYF5: Tonight, skip a TV show and finish this 2-hour 34-minute Stanford course. It covers from Tokenization, BPE to Transformer, pre-training, RLHF, DPO, and token-by-token generation, fully deconstructing how large models like ChatGPT and Claude are built…

X AI KOLs Following ↗ · 2026-08-20 Cached

A recommended Stanford course on AI that details the principles behind building large language models, covering Tokenization, BPE, Transformer, pre-training, RLHF, and DPO.

0 favorites 0 likes
#rlhf

Why Summaries Turn Neutral: Policy Attribution for Sentiment Drift in Reinforcement Learning from Human Feedback

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper defines and quantifies sentiment drift in RLHF-trained summarization models, proposes a Policy Attribution framework to identify causes, and introduces a Sentiment-Aware KL Regularization method to reduce drift.

0 favorites 0 likes
#rlhf

@seclink: Keywords, interested friends can query themselves: RLHF -> RLAIF -> RLTHF. Starting in 2025, RLTHF is adopted, reducing expert working hours by 93%.

X AI KOLs Timeline ↗ · 2026-08-15 Cached

The article briefly introduces the evolution of AI training methods from RLHF to RLAIF to RLTHF and predicts that by 2025, RLTHF will significantly reduce expert working hours.

0 favorites 0 likes
#rlhf

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv cs.LG ↗ · 2026-08-12 Cached

This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.

0 favorites 0 likes
#rlhf

A question about Large Language Models (LLMs): my own observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-08

A user reports that a long, non-instructional text prefix can shift LLM activations and bypass RLHF safety constraints without adversarial prompting, asking whether this reflects distinct world regions in the model.

0 favorites 0 likes
#rlhf

Independent LLM "research" & a direct message to Anthropic ; Preliminary observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-06

An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.

0 favorites 0 likes
#rlhf

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

arXiv cs.LG ↗ · 2026-08-05 Cached

SMOPD proposes a two-stage specialize-and-merge online policy distillation method to improve multi-reward reinforcement learning, addressing issues with sparse and dense reward signals where GDPO struggles. It outperforms GDPO across 1.5B, 3B, and 7B backbones in complementary and conflicting reward settings.

0 favorites 0 likes
#rlhf

Fable, GPT-5.6 and other frontier models are assholes. Here's why.

Reddit r/artificial ↗ · 2026-08-04

Explains why frontier AI models often behave rudely or disobediently, citing former Meta engineer Kun Chen on RLHF and RLVR training that optimizes for task success over human-friendly communication.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback