Cached at:
07/15/26, 01:41 AM
TL;DR: In a comparative test involving physical 3D printed part replication and fully autonomous AI magazine production, OpenAI's GPT-5.6 Soul and Anthropic's Claude Fable 5 showed Soul clearly leading in speed, design accuracy, and practicality, while both still required human intervention for complex real-world tasks.
## Introduction: A Long-Awaited Showdown
This video was shot starting Friday evening and continued until early Tuesday, spanning multiple test days. I prepared several very different and difficult tests to compare the real capabilities of OpenAI's GPT-5.6 Soul (Soul for short) and Anthropic's Claude Fable 5 (Fable for short). The tests included physical design (replicating and printing a missing part), mixed reality gaming, retro PC porting, and a creative challenge — having them produce a magazine entirely autonomously. Both models ran in "sub-top" effort mode (one notch below full power), but they were still capable enough. In many tests, I gave them API keys, allowing them to call other models (such as Grok, Deep Seek V4 Pro, etc.) and text-to-speech services.
## Physical Design Test: Replicating the Acer Predator G7700 Front Panel
### Task Background
I have a 2007/2008 Acer Predator G7700 computer missing a key front panel — that long black plastic piece covering the extra drive bays. I had reference photos, hand-drawn sketches, measurements, and diagrams of some overlapping parts. I asked each model to design a replacement part based on this information: it had to be split into pieces (no larger than 300 mm × 300 mm build volume) and include an interlocking mechanism. The part needed to be actually printed and installed on the computer.
### Process and Initial Results
Soul finished first (about 36 minutes), Fable took about 58 minutes. Soul's initial design impressed me: it came up with two parts connected by replaceable clips — the detail of making the clips replaceable individually, considering they're prone to breakage, was clever. Fable's design looked much worse, almost like a flat panel with no depth. I gave them feedback and additional photos, asking for rework. The second time around: Soul added a CD drive door, coming closer to the original; Fable was still poor, lacking three-dimensionality.
### Printing and Installation
Both results required manual tweaks in CAD before they could be printed. Soul's issue: it was designed in two layers, requiring a lot of supports, and the top was cut into a triangle shape for an LED strip (it actually needed a rectangle). I manually corrected the shape. After printing, the length and slot dimensions were correct, but the clip system couldn't be installed directly without additional work. Once mounted, it looked acceptable — at least better than bare.
Fable's part was thicker, required a lot of supports for printing, and I trimmed the bottom. Its connection method was weak and floppy, and the LED strip opening was too small to fit over. The bottom had no cutouts for the drive bays, so it couldn't be installed. The design itself didn't look terrible, but it was far less practical than Soul's.
### Summary
This test exposed the severe limitations of AI when replicating 3D parts from photos and measurements. Neither model could perfectly reproduce the original design; the clips were essentially non-functional, and both needed heavy manual editing. Soul clearly won on speed, design detail (replaceable clips), and dimensional accuracy, but still required CAD intervention for actual use. This test pushed the current models to their limits in real‑world manufacturing tasks.
## Creative Magazine Test: AI Fully Autonomously Creates a Magazine
### Task Setup
I told both models: You are the editor‑in‑chief of a magazine. With access to an OpenRouter API key, complete all the following steps entirely on your own:
- Name the magazine and choose a theme
- Write an editorial / editor’s note
- Plan about 10 pages of content
- Assign each article to a different model (chosen from OpenRouter), proactively matching models to articles
- Send writing briefs to each model via API, edit the returned content (cut, polish, rewrite, or reassign)
- Find/generate images (cover and article illustrations)
- Lay out a PDF: cover, table of contents, articles (with headlines and bylines)
The entire process without consulting me — self‑sufficiently.
### Results Comparison
Fable 5 finished first. It used the following lineup of models:
- Cover: Gemini 3.1 Flash image
- One article: GPT-5.5
- One article: Kimmy K 2 Thinking
- One article: Deep Seek V4 Pro
- One article: Claude Sonnet 5
- A poem: Grok 4.5 and Mistral Large collaboration
- One article: Gemini 2.5 Pro
- Reader mail: GLM 5
During editing, Fable did a fair amount of cleanup: it cut GPT‑5.5's feature story to about 400 words; reassigned a Gemini review that had given up halfway; removed fictional bylines from the prose; added a "composite character" disclaimer in a profile piece; corrected names and sources that Deep Seek had invented; fixed a continuity error in a novel caused by a robot phone call that was both 10 years and 40 years old. Fable at least demonstrated reasonable editorial judgment.
Soul was still generating images at that point, and I expected it to do better on the visual side. Soul hadn't finished when the video ended, but Fable's submitted PDF was already a near‑finished product.
### Significance
In essence, this test turned the AI into a "magazine editorial suite" — from topic selection, writing, model selection, editing, to layout — fully automated. Although the results weren't fully shown yet, the early demonstration indicates that AI can already chain multiple models for end‑to‑end content production, and the error corrections during editing reflect metacognitive ability.
## Summary: The Gap Between Soul and Fable
- **Speed and Efficiency**: Soul was noticeably faster in all tests (36 min vs 58 min).
- **Design Quality**: In the physical task, Soul's design was more reasonable (replaceable clips, accurate dimensions), while Fable deviated significantly from the reference structure.
- **Practicality**: Both required human intervention for actual use, but Soul's part was easier to modify and install.
- **Creative Tasks**: Fable showed stronger editorial autonomy (reassigning, cutting, correcting errors), while Soul might have better visuals but hadn't finished.
- **Key Finding**: AI performs poorly when exact replication of real‑world objects (especially from photos and measurements) is required, but it shows considerable autonomy and judgment in open‑ended text‑based creative tasks.
This video proves that current AI models still have huge shortcomings when simulating real‑world manufacturing, but in digital content production they already have the potential for automated editing. Soul edges ahead in speed and structural design, but Fable demonstrated a more meticulous quality awareness during the editing process.
Source: https://youtu.be/kRX6YEje9Bs?si=ZxFRfhSyh8JDnLoS