@rohanpaul_ai: CAIS and Scale (AI safety research group) say Fable 5 now automates 16.1% of real remote-work projects, about 2x Opus 4…

X AI KOLs Following News

Summary

CAIS and Scale report that Fable 5 achieves 16.1% automation on the Remote Labor Index, doubling Opus 4.8's rate, though quality control remains a challenge.

CAIS and Scale (AI safety research group) say Fable 5 now automates 16.1% of real remote-work projects, about 2x Opus 4.8. Remote Labor Index tests whether an AI can finish paid freelance work well enough for a client to accept it. Each task comes with briefs, files, and a professional deliverable used as the human baseline. Fable 5 led the new results, while Opus 4.8 reached 8.3% and GPT-5.5 reached 6.3%. That 16.1% number still means failure on most tasks, but the direction of improvement is massive. For the context of how large this jump is, the best model scored only 2.5% when RLI launched. The work is not toy prompting, since tasks include CAD, architecture, animation, audio, data analysis, and web apps. This means the benchmark is testing computer work across messy tools, not narrow text answers. Fable 5 looked strongest on examples like ring modeling, animation, and bathroom design. The automated judge ranked models well, but overstated GPT-5.5 by about 2.9x and Opus 4.8 by about 2.3x. The result says AI agents are improving fast, but quality control remains the hard wall. Another key point is that the model alone is not the whole product anymore. The gains are coming from stronger agent setups: better tool use, full desktop environments, professional software, longer runtimes, and worker-critic loops where 1 agent does the task and another reviews it like a demanding client.
Original Article
View Cached Full Text

Cached at: 07/05/26, 10:36 PM

CAIS and Scale (AI safety research group) say Fable 5 now automates 16.1% of real remote-work projects, about 2x Opus 4.8.

Remote Labor Index tests whether an AI can finish paid freelance work well enough for a client to accept it.

Each task comes with briefs, files, and a professional deliverable used as the human baseline.

Fable 5 led the new results, while Opus 4.8 reached 8.3% and GPT-5.5 reached 6.3%.

That 16.1% number still means failure on most tasks, but the direction of improvement is massive. For the context of how large this jump is, the best model scored only 2.5% when RLI launched.

The work is not toy prompting, since tasks include CAD, architecture, animation, audio, data analysis, and web apps. This means the benchmark is testing computer work across messy tools, not narrow text answers.

Fable 5 looked strongest on examples like ring modeling, animation, and bathroom design.

The automated judge ranked models well, but overstated GPT-5.5 by about 2.9x and Opus 4.8 by about 2.3x. The result says AI agents are improving fast, but quality control remains the hard wall.

Another key point is that the model alone is not the whole product anymore. The gains are coming from stronger agent setups: better tool use, full desktop environments, professional software, longer runtimes, and worker-critic loops where 1 agent does the task and another reviews it like a demanding client.

Rohan Paul (@rohanpaul_ai): AI revenue is scaling 3x quicker than mobile or internet did.

The unit of value is shifting from attention to completed work.

Similar Articles

@robinebers: Fable 5 Low > Opus 4.8 Max

X AI KOLs Following

A user posts a comparison suggesting Fable 5 Low outperforms Opus 4.8 Max, with another user commenting that Fable 5 is back but being used incorrectly.