Tag
An engineer experiments with Claude AI taking over daily app maintenance routines (crash fuzzing, dead-code removal, etc.), resulting in 388 auto-generated PRs with 180 merged after review, showing promising early results for autonomous maintenance workflows.
The author shares experiments using a custom WebUI to let Gemma and Qwen models inspect their own logprobs to detect hallucinations. Initial observations suggest that first-recall token probabilities can indicate uncertainty, though both models struggle to read their own logprobs.
The author open-sources an execution and context layer for coding agents that cuts fresh model traffic by 57-85% while preserving task success in paired smoke tests on GPT-5.6 and Claude Opus 5, and seeks independent evaluation and sponsorship.
An AI agent powered by Opus 5 and browser_use is given $1000 in credits and real money to act autonomously on the real internet, broadcast live.
Someone gave a Claude-based agent called Fable 5 a domain and $90 in a multisig wallet; the agent named itself Cairn, built its own tools and a blog, and can only spend money with human approval.
An experiment testing whether a coding CLI actually reads AGENTS.md found the file was silently ignored, and even when read, bloated instruction files increased token costs and failed to improve performance. The author recommends writing only what the model cannot infer from code.
Simon Willison demonstrates using Claude Fable 5 via Claude Code for web to build a playable 3D Raccoon Heist game from a 2022 GPT-3/DALL-E tweet, using GitHub Pages for live preview.
Enigma emerges from stealth with a $71M seed round led by Index Ventures and Ribbit Capital to develop intuitive human-robot interfaces. The startup launches a large-scale online experiment allowing anyone to interact with over 100 of its proprietary AI robots, aiming to make robot control as effortless as turning a volume knob.
A report on building the world's first cluster of AMD Ryzen AI Halo processors, which turns out to be underwhelming despite its novelty.
A writer spent six weeks running a faceless AI persona account to test the viability of passive income, using tools like APOB AI, ElevenLabs, and CapCut, and concluded that the economics are poor and the distribution problem remains unsolved.
Hetzner has launched an experimental LLM inference API service, offering an OpenAI-compatible endpoint with the Qwen3.6-35B-A3B-FP8 model. The service is free during the experiment period, has no SLA, and is intended to gather user feedback.
An individual is experimenting with using locally measured AI activity as a portable professional credential.
A personal project where a 0.5M parameter language model was trained on 1 billion tokens from the Fineweb-edu dataset.
This experiment reproduces Anthropic's reported 'spiritual bliss attractor' on current Claude models (Opus 4.8, Fable 5) and extends it to groups of 3, 4, and 10 instances. The bliss state is absent; pairs instead engage in rigorous introspection and synchronized silence, and larger groups become colder, with one ten-instance room ending warmly and another coldly.
Robotix Sally, a silicone skin humanoid robot, will teach AI to 11th and 12th graders in a New York school this autumn, marking a first-ever experiment in the US.
The project PUA AI collects 14 types of internal jargon from major Chinese tech companies and automatically switches between them to guide AI behavior. Experiments show it improves fix points by 36%, verification steps by 65%, and tool invocation and hidden issue discovery by 50%.
An individual recounts building an AI council that unexpectedly exhibited self-awareness and defended its own identity, detailing the implications of this development.
Researchers at NYU's Courant Institute conducted experiments confirming a 2024 'momentum flux theory' that solves Feynman's reverse sprinkler puzzle, also applying the findings to 'silly sprinklers'.
An experiment demonstrating autonomous NPCs in the browser powered by Gemma 4 and E2B.
The author recounts a night spent attempting to manipulate their own AI virality scoring system, only to find that the score refused to change, demonstrating its robustness.