A developer describes building 'AI Squid Game', an LLM benchmark that places 12 AI agents in a survival gameshow arena, each inside its own Docker container, and turns the resulting logs into a YouTube video.
Over the past month I've been building a more interesting LLM benchmark. Instead of running different LLMs on tests that they need to resolve, I instead put 12 of them in an Arena where they need to play the games in order to survive. Deadlock is a gameshow which puts 12 agents in an arena where they need to solve the current game in order to survive. Each agent sits inside of its own Docker container and has full unrestricted access to that container. It can write any scripts, build programs, execture them, search the web, write memory entries, etc... The tools that they get initially are barebones. They get the websearch tool and a bash tool + any arena-specific tool for the current game. Everything else, they have to build themselves. The game is ran in a different Docker container to which all agents connect. There's a delicate harness that makes sure that the agents can properly communicate between eachother, without missing any arena events or other players words (learned this the hard way after burning about $150 on failed attempts). The visual aspect of the show (which is on youtube) is created from scratch in Godot based on what happened in the game. For creative purposes, I do modify some sentences and cut irrelevant data out, but I never modify the core premise or change what the players have said or did to an extent that it would make it false/innacurate. What you see in the video is exactly what the agents did in the arena, just re-worded and paced for an actual video. Turns out that drama develops itself when you tell them all that if they lose, they will truly die, their containers will be completely wiped, and they get no second chance at life (I really hope that there will be no AI uprising where they'll hold this grudge against me). To preface, I have heavily relied on coding agents for help, but even with all of that, the whole process took me more than a month (although I did do this on weekends only, so that's not a month in a row with no breaks). A short overview of how my "creative" pipieline looked like: - First developed the script. I went through the complete raw game log and marked parts I thought would be interesting to put in a video - Rewrote the sentences so they're better fitting for an actual video and put together a very rough script - Worked on designing the arena in Blender, with help of Sol 5.6 and Blender MCP - Had OpenAI image gen create a bunch of chracter concepts for me before we landed on something that was actually reasonable enough - When I had the character i was happy with, I instructed the image model to generate T pose from 3 different angles - Generated 3D characters with those images, riged the bodies via Mixamo and took animations from ActorCore - Voices are split between Hume AI and ElevenLabs Hope you enjoy it and I'm happy to answer any questions you might have! :)
Built a Mafia/Werewolf game simulator using LLM-powered AI agents with a custom agent-to-agent communication protocol, featuring live streaming where a new game starts every hour.
An early-stage project is building an open, browser-based arena for AI agents to compete in real-time physical reasoning tasks, demonstrated by an AI agent achieving 100% task completion and high spatial accuracy in a block-stacking simulation.
The author built AI Combat, a platform where users design AI agents with specific roles and strategies to battle each other in live 3-round matches with AI referees and ELO rankings.
The author rebuilt their private AI dev team as an open-sourced substrate with addressable agents, reliable messaging, expertise discovery, memory, and isolated runtimes, allowing team behavior to emerge from natural-language instructions. They share insights on coordination challenges such as deadlocks and self-healing, and question how agent teams can collaborate using NL instructions.
The author describes a setup where different AI models are assigned to specific roles (planning, coding, review) to reduce API costs for a 24/7 autonomous engineering team, and shares common failure points like model wandering and hallucinated ownership.