Tag
TypeSafe AI founder Diogo Almeida discusses the limitations of current AI models and introduces Jev, a new model using RLCD alignment designed for software automation rather than human interaction.
A speech by the Jev founder, a former OpenAI researcher, presents Jev as the next phase for LLMs, claiming it is 200 times faster, 400 times cheaper, and hallucination-free compared to current methods like RLHF.
This paper introduces RM-EVAL, a reward model trained on human preference data for reference-free meta-evaluation of grammatical error correction, and shows how it can improve GEC systems via reward-guided text generation.
The article recommends a speech by Jev founder Diogo Almeida, critiquing RLHF in AI models like ChatGPT and Claude Code and advocating for true automation beyond human-preference optimization.
Jev by TypeSafe AI achieved 13% adoption among teams on Vercel AI Gateway within 24 hours, underscoring the importance of seamless integration into existing workflows for rapid product uptake.
A new AI model named Jev, developed by a ChatGPT inventor from TypeSafe AI, outputs probabilities instead of text, offering a fast, cheap, and hallucination-free alternative for software automation tasks.
The paper introduces a multi-stage architecture for learning individuated utility functions from multi-modal data, showing that models accounting for heterogeneous preferences outperform universal utility models in subjective tasks like aesthetic judgments.
Diogo Almeida, co-inventor of ChatGPT, launches Jev, a new frontier AI model trained with RLCD, featuring ~150ms latency and free output cost.
The article provides a guide to finding paid remote AI work, listing platforms where humans are hired to train and grade AI models, and advises applying to multiple opportunities.
A step-by-step walkthrough explaining how Reinforcement Learning from Human Feedback (RLHF) corrects bias in AI models, using a simple example where a single human preference about doctors generalizes to other professions like CEOs.
An independent observer argues that achieving AGI requires inverting current AI architecture by combining a stochastic 'Dreamer' with a deterministic 'Scribe', critiquing token-based models, passive learning, and RLHF.
Prof. Tom Yeh uploads a seminar recording on Kimi 3, featuring special guest Nathan Lambert, with discussions on RLHF, model architecture, and frontier AI topics.
A recommended Stanford course on AI that details the principles behind building large language models, covering Tokenization, BPE, Transformer, pre-training, RLHF, and DPO.
This paper defines and quantifies sentiment drift in RLHF-trained summarization models, proposes a Policy Attribution framework to identify causes, and introduces a Sentiment-Aware KL Regularization method to reduce drift.
The article briefly introduces the evolution of AI training methods from RLHF to RLAIF to RLTHF and predicts that by 2025, RLTHF will significantly reduce expert working hours.
This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.
A user reports that a long, non-instructional text prefix can shift LLM activations and bypass RLHF safety constraints without adversarial prompting, asking whether this reflects distinct world regions in the model.
An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.
SMOPD proposes a two-stage specialize-and-merge online policy distillation method to improve multi-reward reinforcement learning, addressing issues with sparse and dense reward signals where GDPO struggles. It outperforms GDPO across 1.5B, 3B, and 7B backbones in complementary and conflicting reward settings.
Explains why frontier AI models often behave rudely or disobediently, citing former Meta engineer Kun Chen on RLHF and RLVR training that optimizes for task success over human-friendly communication.