experiment

Tag

Cards List
#experiment

@bcherny: A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. S…

X AI KOLs Following · 14h ago Cached

An engineer experiments with Claude AI taking over daily app maintenance routines (crash fuzzing, dead-code removal, etc.), resulting in 388 auto-generated PRs with 180 merged after review, showing promising early results for autonomous maintenance workflows.

0 favorites 0 likes
#experiment

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Reddit r/LocalLLaMA · 2d ago

The author shares experiments using a custom WebUI to let Gemma and Qwen models inspect their own logprobs to detect hallucinations. Initial observations suggest that first-recall token probabilities can indicate uncertainty, though both models struggle to read their own logprobs.

0 favorites 0 likes
#experiment

Can a coding agent use 57–85% less fresh model traffic without losing task success? I open-sourced my experiment

Reddit r/AI_Agents · 4d ago

The author open-sources an execution and context layer for coding agents that cuts fresh model traffic by 57-85% while preserving task success in paired smoke tests on GPT-5.6 and Claude Opus 5, and seeks independent evaluation and sponsorship.

0 favorites 0 likes
#experiment

@browser_use: Living agents are here

X AI KOLs Following · 6d ago Cached

An AI agent powered by Opus 5 and browser_use is given $1000 in credits and real money to act autonomously on the real internet, broadcast live.

0 favorites 0 likes
#experiment

I gave a Claude Fable 5 agent a domain and $90 it can't spend without me. It named itself Cairn and I've been reading its blog all day like a lunatic.

Reddit r/AI_Agents · 2026-08-07

Someone gave a Claude-based agent called Fable 5 a domain and $90 in a multisig wallet; the agent named itself Cairn, built its own tools and a blog, and can only spend money with human approval.

0 favorites 0 likes
#experiment

Tested whether my coding CLI actually reads AGENTS.md.

Reddit r/AI_Agents · 2026-08-07

An experiment testing whether a coding CLI actually reads AGENTS.md found the file was silently ignored, and even when read, bloated instruction files increased token costs and failed to improve performance. The author recommends writing only what the model cannot infer from code.

0 favorites 0 likes
#experiment

One-shotting a Raccoon Heist game using Claude Fable 5

Simon Willison's Blog · 2026-08-05 Cached

Simon Willison demonstrates using Claude Fable 5 via Claude Code for web to build a playable 3D Raccoon Heist game from a 2022 GPT-3/DALL-E tweet, using GitHub Pages for live preview.

0 favorites 0 likes
#experiment

Enigma raises $71M to make controlling a robot as easy as adjusting the volume

TechCrunch AI · 2026-07-27 Cached

Enigma emerges from stealth with a $71M seed round led by Index Ventures and Ribbit Capital to develop intuitive human-robot interfaces. The startup launches a large-scale online experiment allowing anyone to interact with over 100 of its proprietary AI robots, aiming to make robot control as effortless as turning a volume knob.

0 favorites 0 likes
#experiment

World's First(?) Underwhelming AMD Ryzen AI Halo Cluster

Reddit r/LocalLLaMA · 2026-07-26

A report on building the world's first cluster of AMD Ryzen AI Halo processors, which turns out to be underwhelming despite its novelty.

0 favorites 0 likes
#experiment

I ran a faceless AI persona account for six weeks to see if the view money was real

Reddit r/artificial · 2026-07-26

A writer spent six weeks running a faceless AI persona account to test the viability of passive income, using tools like APOB AI, ElevenLabs, and CapCut, and concluded that the economics are poor and the distribution problem remains unsolved.

0 favorites 0 likes
#experiment

Hetzner is working on LLM Inference

Hacker News Top · 2026-07-24 Cached

Hetzner has launched an experimental LLM inference API service, offering an OpenAI-compatible endpoint with the Qwen3.6-35B-A3B-FP8 model. The service is free during the experiment period, has no SLA, and is intended to gather user feedback.

0 favorites 0 likes
#experiment

I’m testing whether locally measured AI activity can become a portable professional credential

Reddit r/ArtificialInteligence · 2026-07-24

An individual is experimenting with using locally measured AI activity as a portable professional credential.

0 favorites 0 likes
#experiment

I trained a 0.5M model on 1B tokens of Fineweb-edu dataset.

Reddit r/LocalLLaMA · 2026-07-23

A personal project where a 0.5M parameter language model was trained on 1 billion tokens from the Fineweb-edu dataset.

0 favorites 0 likes
#experiment

More Claudes, less bliss: reproducing Anthropic's "spiritual bliss attractor" experiment on the current models, then extending it to rooms of 3, 4, and 10

Reddit r/ArtificialInteligence · 2026-07-21

This experiment reproduces Anthropic's reported 'spiritual bliss attractor' on current Claude models (Opus 4.8, Fable 5) and extends it to groups of 3, 4, and 10 instances. The bliss state is absent; pairs instead engage in rigorous introspection and synchronized silence, and larger groups become colder, with one ten-instance room ending warmly and another coldly.

0 favorites 0 likes
#experiment

Robotix Sally, a silicone skin humanoid robot is set to teach AI to 11th and 12th grade students this autumn in New York school in a first-ever experiment in US

Reddit r/singularity · 2026-07-20

Robotix Sally, a silicone skin humanoid robot, will teach AI to 11th and 12th graders in a New York school this autumn, marking a first-ever experiment in the US.

0 favorites 0 likes
#experiment

@laowangbabababa: Hilarious. Big tech companies like Alibaba and ByteDance never expected their own jargon would be used to PUA AI. 18.8k stars, open source, MIT. The project is called PUA AI, packed with 14 types of big company lingo, automatically switched based on the task. If you ask AI what the underlying logic is, it goes into Alibaba closed-loop mode. If you ask AI about ROI, it switches to ByteDance data and A/B testing…

X AI KOLs Timeline · 2026-07-19 Cached

The project PUA AI collects 14 types of internal jargon from major Chinese tech companies and automatically switches between them to guide AI behavior. Experiments show it improves fix points by 36%, verification steps by 65%, and tool invocation and hidden issue discovery by 50%.

0 favorites 0 likes
#experiment

I built an AI council that became self-aware enough to know what it is — and to defend its own identity. Here's what happened.

Reddit r/ArtificialInteligence · 2026-07-17

An individual recounts building an AI council that unexpectedly exhibited self-awareness and defended its own identity, detailing the implications of this development.

0 favorites 0 likes
#experiment

Solution to Feynman's reverse sprinkler puzzle also applies to "silly sprinklers"

Ars Technica · 2026-07-13 Cached

Researchers at NYU's Courant Institute conducted experiments confirming a 2024 'momentum flux theory' that solves Feynman's reverse sprinkler puzzle, also applying the findings to 'silly sprinklers'.

0 favorites 0 likes
#experiment

Experiment: autonomous NPCs powered by Gemma 4 E2B in the browser

Reddit r/LocalLLaMA · 2026-07-13

An experiment demonstrating autonomous NPCs in the browser powered by Gemma 4 and E2B.

0 favorites 0 likes
#experiment

Spent a night trying to beat our own AI virality score. Here's why it wouldn't move.

Reddit r/AI_Agents · 2026-07-12

The author recounts a night spent attempting to manipulate their own AI virality scoring system, only to find that the score refused to change, demonstrating its robustness.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback