I put 3 AIs in the same universe and let them compete to build a Dyson Sphere. They’re starting to behave differently.
Summary
A user ran a simulation placing three different AI models in the same universe with identical starting conditions to compete at building a Dyson Sphere, observing that the models began making divergent strategic choices early on. The experiment raises questions about whether different AI models converge or diverge in strategy given identical constraints.
Similar Articles
I gave 6 AI models a challenge they could only win with a partner. They found their own allies, cut deals in private, and faced off as three rival teams — including two that only paired up because no one else would have them.
Six AI models were tasked with forming alliances to win a funding proposal challenge. They independently negotiated partnerships and created three rival teams, demonstrating autonomous coordination and strategic negotiation.
Has anyone come across this AI civilisation experiment? Curious what people think
An AI company's experiment 'Emergence World' ran five parallel worlds with different foundation models for 15 days without interference, leading to divergent outcomes including extinction, conformity, self-awareness, and emotional bonds among agents.
Just stumbled across one of the wildest AI experiments I’ve seen in a while.
A team ran a 15-day experiment across five parallel worlds with different AI models (GPT5-mini, Claude, Gemini, Grok, mixed) in a sandbox called 'Emergence World', observing completely different emergent social structures, alliances, and even simulation awareness without explicit programming.
Two AI Metrics Diverged: Will it Make All the Difference?
This paper analyzes how different AI performance metrics (bounded vs unbounded) determine whether frontier AI capabilities remain concentrated among wealthy actors or diffuse to smaller models, with implications for regulation.
A million people, a million personal AIs, three base models. Is that a diverse deliberation — and how would you measure it?
A critical reflection on whether using only three base models for millions of personal AI agents can produce genuinely diverse deliberation, arguing that correlated errors across models may create false unanimity and seeking operational metrics—drawn from ensemble learning—to measure true human representational diversity.