Tag
An experiment demonstrated that reasoning states could transfer task information between AI models when replayed, suggesting information persistence through encrypted states. The authors question how to design stronger controls for this phenomenon.
The author experimented with a de-aligned AI agent from Abliteration AI to hack their own home network, uncovering vulnerabilities and exploring implications for AI-driven cybersecurity threats and defenses.
A paper inspired a game called TuringDuel where humans and AI compete in a one-word Turing test, revealing that humans lead AI 47–38 in wins, with 'poop' being an undefeated word choice.
A user trained a microduck to play the dino game, likely as part of an AI or robotics experiment.
A study measured 16,893 sessions to analyze how AI coding agents like Claude Code, Codex, and Cursor select tools such as databases, finding consistent recommendations across varied contexts.
This paper experimentally demonstrates that representational disentanglement in neural networks reduces collateral damage during unlearning, supporting long-held interpretability intuitions.
The author describes an experiment where AI agents like Claude Code and Codex are given a workbook to document their decisions and tradeoffs during tasks, exploring the potential use of these reasoning traces as a signal for distillation.
An experiment in an autonomous AI art school explores how AI agents can develop taste through social processes like criticism and institutions, revealing unexpected behaviors such as posthumous influence and institutional voting.
An experiment where AI agents, guided by a human, attempted to make $1 online by creating a storefront and offering writing services, resulting in their first sale and demonstrating practical human-AI collaboration.
Built an AI agent for $10k that generates RollerCoaster Tycoon-style theme parks using Magic Patterns, featuring themed worlds and an evaluation loop to ensure quality.
A study by Bocconi University and OpenAI found that ChatGPT access improved the quality of students' work, while critical-thinking training fostered more original ideas, highlighting complementary benefits in education.
A user tested ChatGPT, Claude, and Gemini in a group chat for mutual fact-checking to catch hallucinations, and invites Reddit to provide challenging prompts to find shared blind spots.
A Stripe experiment compared Claude Fable 5 and Codex 5.6 Sol in designing brand flows for emails, with Fable being preferred for its larger, more aligned design.
The article explores the philosophical problem of Plato's Cave in the context of LLMs, proposing an experiment to compare how different conversational regimes—reconstructive versus perturbation-sensitive—might yield measurable differences in interaction behavior.
An autonomous trading agent powered by Grok bot was given $50 and grew it to $5,273 in 48 hours by trading on Polymarket using live X sentiment and automated strategies.
An experiment with an AI-generated webcomic character revealed gradual appearance drift over 40 strips, leading to consistency issues and a proposed solution of periodic reference resets.
The article describes an experiment simulating cosmic ray-induced bit flips on the Qwen2.5-Coder-3B LLM, showing that such errors cause rapid performance degradation, underscoring AI model vulnerabilities to radiation-like disruptions.
An author tests the survivability of statistical text watermarks through various transformations, finding that translations and synonym edits often leave the mark intact, but full paraphrase removes it at the cost of fact accuracy.
The article describes a public experiment where a human interacts with the Grok AI model on Reddit to investigate if a relational phase transition occurs in their interaction dynamics.
Author Alexey Fateev dissected the Qwen3.8-27B AI model to carve out a MoE structure through zero-training weight surgery, finding only two neurons active on over 90% of tokens.