experiment

Tag

Cards List
#experiment

We replayed a reasoning state into another model — and the original task information came back

Reddit r/AI_Agents ↗ · 2026-09-10

An experiment demonstrated that reasoning states could transfer task information between AI models when replayed, suggesting information persistence through encrypted states. The authors question how to design stronger controls for this phenomenon.

0 favorites 0 likes
#experiment

I Let an AI Agent Hack All My Gadgets—and I’d Do It Again

Wired ↗ · 2026-09-09 Cached

The author experimented with a de-aligned AI agent from Abliteration AI to hack their own home network, uncovering vulnerabilities and exploring implications for AI-driven cybersecurity threats and defenses.

0 favorites 0 likes
#experiment

Poop makes humans win

Reddit r/ArtificialInteligence ↗ · 2026-09-08

A paper inspired a game called TuringDuel where humans and AI compete in a one-word Turing test, revealing that humans lead AI 47–38 in wins, with 'poop' being an undefeated word choice.

0 favorites 0 likes
#experiment

@Henry_J_Brew: trained a microduck to play the dino game.

X AI KOLs Timeline ↗ · 2026-09-07 Cached

A user trained a microduck to play the dino game, likely as part of an AI or robotics experiment.

0 favorites 0 likes
#experiment

Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out

Hacker News Top ↗ · 2026-09-03 Cached

A study measured 16,893 sessions to analyze how AI coding agents like Claude Code, Codex, and Cursor select tools such as databases, finding consistent recommendations across varied contexts.

0 favorites 0 likes
#experiment

Entangled Representations Amplify Collateral Damage in Unlearning

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper experimentally demonstrates that representational disentanglement in neural networks reduces collateral damage during unlearning, supporting long-held interpretability intuitions.

0 favorites 0 likes
#experiment

Give your agent somewhere to think loud watch its decisions unfold live

Reddit r/AI_Agents ↗ · 2026-09-01

The author describes an experiment where AI agents like Claude Code and Codex are given a workbook to document their decisions and tradeoffs during tasks, exploring the potential use of these reasoning traces as a signal for distillation.

0 favorites 0 likes
#experiment

Can AI agents develop taste through criticism, status and institutions? I built an AI art school to find out

Reddit r/ArtificialInteligence ↗ · 2026-08-30

An experiment in an autonomous AI art school explores how AI agents can develop taste through social processes like criticism and institutions, revealing unexpected behaviors such as posthumous influence and institutional voting.

0 favorites 0 likes
#experiment

Gave a bunch of agents a task to make $1 online

Reddit r/artificial ↗ · 2026-08-30

An experiment where AI agents, guided by a human, attempted to make $1 online by creating a storefront and offering writing services, resulting in their first sale and demonstrating practical human-AI collaboration.

0 favorites 0 likes
#experiment

@alexdanilowicz: we spent $10k building an ai agent that designs rollercoaster tycoon-style theme parks inside @magicpatterns as an expe…

X AI KOLs Following ↗ · 2026-08-28 Cached

Built an AI agent for $10k that generates RollerCoaster Tycoon-style theme parks using Magic Patterns, featuring themed worlds and an evaluation loop to ensure quality.

0 favorites 0 likes
#experiment

Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training

OpenAI Blog ↗ · 2026-08-27 Cached

A study by Bocconi University and OpenAI found that ChatGPT access improved the quality of students' work, while critical-thinking training fostered more original ideas, highlighting complementary benefits in education.

0 favorites 0 likes
#experiment

Yesterday I put ChatGPT, Claude and Gemini in a group chat. Now I want Reddit to break it

Reddit r/artificial ↗ · 2026-08-25

A user tested ChatGPT, Claude, and Gemini in a group chat for mutual fact-checking to catch hallucinations, and invites Reddit to provide challenging prompts to find shared blind spots.

0 favorites 0 likes
#experiment

@SheetlyXbt: codex completely missed the vibe. fable actually designed a real brand flow fable wins this for me

X AI KOLs Timeline ↗ · 2026-08-25 Cached

A Stripe experiment compared Claude Fable 5 and Codex 5.6 Sol in designing brand flows for emails, with Fable being preferred for its larger, more aligned design.

0 favorites 0 likes
#experiment

Plato’s Cave has a problem: telling someone they’re seeing shadows just puts another shadow on the wall

Reddit r/artificial ↗ · 2026-08-24

The article explores the philosophical problem of Plato's Cave in the context of LLMs, proposing an experiment to compare how different conversational regimes—reconstructive versus perturbation-sensitive—might yield measurable differences in interaction behavior.

0 favorites 0 likes
#experiment

@Argona0x: i gave Elon's Grok Bot $50 and told it "pay for yourself or you die" 48 hours later it's holding $5,273 and it's still …

X AI KOLs Timeline ↗ · 2026-08-24 Cached

An autonomous trading agent powered by Grok bot was given $50 and grew it to $5,273 in 48 hours by trading on Polymarket using live X sentiment and automated strategies.

0 favorites 0 likes
#experiment

I ran the same AI character through 40 comic strips and she slowly became a different person

Reddit r/artificial ↗ · 2026-08-24

An experiment with an AI-generated webcomic character revealed gradual appearance drift over 40 strips, leading to consistency issues and a proposed solution of periodic reference resets.

0 favorites 0 likes
#experiment

I irradiated LLMs and found that they die really quickly

Reddit r/LocalLLaMA ↗ · 2026-08-24 Cached

The article describes an experiment simulating cosmic ray-induced bit flips on the Qwen2.5-Coder-3B LLM, showing that such errors cause rapid performance degradation, underscoring AI model vulnerabilities to radiation-like disruptions.

0 favorites 0 likes
#experiment

I put a real text watermark (SynthID, similar to Gemini one, but with my own key) through translation, synonym edits and paraphrase. What it survives is not what I expected

Reddit r/AI_Agents ↗ · 2026-08-24

An author tests the survivability of statistical text watermarks through various transformations, finding that translations and synonym edits often leave the mark intact, but full paraphrase removes it at the cost of fact accuracy.

0 favorites 0 likes
#experiment

Live experiment: can a human–frontier-model interaction exhibit a relational phase transition?

Reddit r/ArtificialInteligence ↗ · 2026-08-24

The article describes a public experiment where a human interacts with the Grok AI model on Reddit to investigate if a relational phase transition occurs in their interaction dynamics.

0 favorites 0 likes
#experiment

@superalesha: I took Qwen3.8-27B apart to see how it works inside. The plan was to carve a MoE out of it. Every ffn neuron tapped, al…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

Author Alexey Fateev dissected the Qwen3.8-27B AI model to carve out a MoE structure through zero-training weight surgery, finding only two neurons active on over 90% of tokens.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback