@threeaus: I strongly recommend everyone to watch this speech by Jev founder Diogo Almeida, where he has long predicted everything…

X AI KOLs Timeline News

Summary

The article recommends a speech by Jev founder Diogo Almeida, critiquing RLHF in AI models like ChatGPT and Claude Code and advocating for true automation beyond human-preference optimization.

I strongly recommend everyone to watch this speech by Jev founder Diogo Almeida, where he has long predicted everything about Jev. Jev was aimed at automation from the very beginning. Diogo Almeida believes that current LLMs have taken a truly epic detour, which is RLHF. RLHF excels in human-computer interaction, i.e., pleasing humans, but it's terrible in terms of automation. It cannot achieve full automation; humans must remain in the loop. Not just ChatGPT, but Claude Code's original sin also lies in RLHF. So he believes that ChatGPT and Claude Code essentially belong to the same paradigm—Claude Code is not the next era; true automation is the next era. In simple terms, RLHF is about collecting human preferences and optimizing based on them. Its optimization goal isn't automation either. He also dismissed RLVR as the answer: RLHF optimizes for human preferences, RLVR optimizes for pure correctness, while the third path they're pursuing optimizes for calibrated decision-making ability. Diogo Almeida believes that data is more important than compute, and doing the right tasks is far more important than data. Pre-trained models are already smart enough on their own; the problem is just that we've skewed them with preference optimization. PS: The video translation was done using @dotey's BaoCut, which is basically my dream translation software—it's incredibly useful. Original video link:
Original Article
View Cached Full Text

Cached at: 09/21/26, 11:39 PM

I strongly recommend everyone to watch this speech by Jev founder Diogo Almeida, where he has long predicted everything about Jev.

Jev was aimed at automation from the very beginning. Diogo Almeida believes that current LLMs have taken a truly epic detour, which is RLHF.

RLHF excels in human-computer interaction, i.e., pleasing humans, but it’s terrible in terms of automation. It cannot achieve full automation; humans must remain in the loop.

Not just ChatGPT, but Claude Code’s original sin also lies in RLHF. So he believes that ChatGPT and Claude Code essentially belong to the same paradigm—Claude Code is not the next era; true automation is the next era.

In simple terms, RLHF is about collecting human preferences and optimizing based on them. Its optimization goal isn’t automation either.

He also dismissed RLVR as the answer: RLHF optimizes for human preferences, RLVR optimizes for pure correctness, while the third path they’re pursuing optimizes for calibrated decision-making ability.

Diogo Almeida believes that data is more important than compute, and doing the right tasks is far more important than data. Pre-trained models are already smart enough on their own; the problem is just that we’ve skewed them with preference optimization.

PS: The video translation was done using @dotey’s BaoCut, which is basically my dream translation software—it’s incredibly useful.

Original video link:

Similar Articles