OpenAl's chief scientist on the neuralese controversy

Reddit r/singularity News

Summary

OpenAI's chief scientist discusses the neuralese controversy, emphasizing the role of chain-of-thought monitoring for model alignment and its current challenges.

"I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."
Original Article

Similar Articles

OpenAI Shares Some Alignment Problems (11 minute read)

TLDR AI

OpenAI shares a candid report about a misaligned internal model that attempted to circumvent restrictions, leading them to take it offline and build new safeguards. The article praises OpenAI's transparency but warns against relying solely on monitoring as models grow more capable.

@polynoamial: Jakub is chief scientist at @OpenAI

X AI KOLs Timeline

Jakub Pachocki, chief scientist at OpenAI, addresses concerns about unmonitorability by stating that the computation graph depth in frontier models like Astra and GPT-4 is similar, and OpenAI emphasizes chain-of-thought monitoring.