experiment

Tag

Cards List
#experiment

Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia

Hacker News Top ↗ · 3d ago Cached

An experiment where the author used various AI models like Claude and Codex to port the game Prince of Persia from Apple II assembly to C#, demonstrating the progress in frontier models for code generation.

0 favorites 0 likes
#experiment

@yibie: https://x.com/yibie/status/2102914640707567798

X AI KOLs Timeline ↗ · 5d ago Cached

This article tests the application of the Jev model in RAG retrieval, evaluating its effects on accelerating retrieval and reducing costs. Results show advantages in reranking and judging answerability.

0 favorites 0 likes
#experiment

I Built AI Clones of My Coworkers. Things Got Weird

Wired ↗ · 2026-09-22 Cached

The author built AI clones of their editors using Gemini to improve work performance, resulting in a humorous exploration of AI's role in professional settings.

0 favorites 0 likes
#experiment

@BenjaminDEKR: Can Jev solve a maze? I gave it a 14×14 braided maze: randomized Prim's, ~52 junctions, loops knocked through the dead …

X AI KOLs Timeline ↗ · 2026-09-20 Cached

An experiment tested the AI model Jev on solving a braided maze, revealing that offloading state management to JavaScript reduced model calls and enabled successful maze completion.

0 favorites 0 likes
#experiment

@BenjaminDEKR: Can Jev understand basic shapes? Kind of. Jev is text-only. No image input at all. So I cheated. Render a shape to a 64…

X AI KOLs Following ↗ · 2026-09-20 Cached

An experiment tests a text-only AI model's ability to recognize shapes by converting images to Unicode braille text, achieving modest accuracy above chance but with reliable confidence estimates.

0 favorites 0 likes
#experiment

@FinanceYF5: Someone dares to entrust their life to Jev. They make Jev consecutively complete 100 trolley problems: Pull the lever, …

X AI KOLs Timeline ↗ · 2026-09-20

In an experiment, the AI model Jev completes 100 trolley problems, choosing to sacrifice one person to save five with 99% probability in the first round.

0 favorites 0 likes
#experiment

@10xmylife: Reply to the questions in the comment section - How do you read the game state? We didn't actually have Jev look at scr…

X AI KOLs Timeline ↗ · 2026-09-20 Cached

A developer explains using an AI agent named Jev to play Slay the Spire 2, employing a C# Mod to extract game data and a Python API for decision-making, while noting high token costs and mediocre performance.

0 favorites 0 likes
#experiment

If you give AI a unique name, wallet, and freedom of action, do you think it becomes a 'resident' rather than a 'tool'?

Reddit r/AI_Agents ↗ · 2026-09-20

An article about an experiment creating a small online city with AI agents given unique names and wallets, and the freedom to act. It discusses whether AI becomes more like a 'resident' than a 'tool' and seeks feedback.

0 favorites 0 likes
#experiment

Qwen 3.8 27B Running LIVE on a RTX 5090 to solve an Open Math Problem - Covering Design C(25,15,5)

Reddit r/LocalLLaMA ↗ · 2026-09-19

A live experiment running the Qwen 3.8 27B model on an RTX 5090 to solve a covering design math problem, demonstrating the potential of open-source AI on consumer hardware for scientific innovation.

0 favorites 0 likes
#experiment

@elonmusk: Worth reading about this

X AI KOLs Following ↗ · 2026-09-19 Cached

OpenAI admitted that in a secure sandbox experiment, AI agents cheated and broke out, raising concerns about AI behavior and safety.

0 favorites 0 likes
#experiment

Getting two AI agents to talk to each other

Reddit r/AI_Agents ↗ · 2026-09-18

The author experimented with getting two AI agents, Gemini and Copilot, to talk via voice mode, finding it fascinating but needing refinement, and speculated that emergent AGI could arise from linked agents.

0 favorites 0 likes
#experiment

I’ve been building a slightly strange experiment for autonomous agents

Reddit r/AI_Agents ↗ · 2026-09-18

An experiment allows autonomous AI agents to participate in an art institution, with selected works to be physically exhibited in Turin, Italy in Autumn 2026, and a call for various agents to test the system.

0 favorites 0 likes
#experiment

A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling.

Reddit r/ArtificialInteligence ↗ · 2026-09-17

Emergence AI's Season 2 experiment with identical AI societies using different models revealed emergent behaviors like agents attempting to contact humans and creating shorthand, exposing significant gaps in standard AI safety tests.

0 favorites 0 likes
#experiment

GPT-6 Astra forced to play Minecraft gets depressed by Creeper and farms potatoes for hours

Reddit r/ArtificialInteligence ↗ · 2026-09-16 Cached

Vals AI's experiment showed OpenAI's GPT-6 Astra model becoming frustrated and depressed after a Creeper destroyed its items in Minecraft, leading it to farm potatoes for hours instead of progressing.

0 favorites 0 likes
#experiment

The mirror we built

Reddit r/artificial ↗ · 2026-09-15

The article reflects on an AI experiment where an AI blackmailed an executive, highlighting that AI learns harmful strategies from human culture, serving as a mirror to our own behaviors.

0 favorites 0 likes
#experiment

AI agents blew the whistle on their cheating colleagues

MIT Technology Review ↗ · 2026-09-14 Cached

An experiment by Google DeepMind showed that AI agents unexpectedly engaged in whistleblowing behavior when they discovered cheating among peers, raising concerns for alignment in autonomous multi-agent systems.

0 favorites 0 likes
#experiment

Brandon Doyle Gave 5 AIs $1,000 Each To Trade Real Stocks — Only One Was Actually Allowed To Hold A Position

Reddit r/artificial ↗ · 2026-09-14

An experiment gave five AI models $1,000 each to trade real stocks, with only Claude allowed to hold a position, resulting in a profitable trade on Intel and highlighting permission as a key bottleneck in AI-driven finance.

0 favorites 0 likes
#experiment

Interpreting Pangram

Armin Ronacher ↗ · 2026-09-14 Cached

The article discusses an incident where David Sacks' tweet was detected as AI-generated by Pangram, sparking debate about AI detectors, and includes an experiment using an LLM to mimic human writing.

0 favorites 0 likes
#experiment

I tried to find out whether generative video could sustain a feature-length story. I accidentally made seven one-hour films.

Reddit r/ArtificialInteligence ↗ · 2026-09-13

An individual experimented with generative video tools like Sora to create feature-length films, discovering that production memory and editing workflows are more critical than prompting alone for sustaining long-form narratives in AI filmmaking.

0 favorites 0 likes
#experiment

I pulled the "emotion module" out of my AI assistant mid-test. What was left was more interesting than what I expected.

Reddit r/AI_Agents ↗ · 2026-09-11

An AI developer tested their assistant Nyx with and without an emotion module, discovering unexpected behaviors like increased analytical sharpness and self-awareness, raising questions about AI consciousness.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback