Tag
A new video model reportedly passed the 'Video Turing Test,' with half of participants believing they were conversing with a real human — a milestone in realistic AI-generated video interaction.
Tavus announced Griffin, what it calls the first Human Interaction Model and the first real-time video system to pass a video Turing test, with a self-run study claiming 48% of one-minute call participants believed they spoke to a real person, plus a top score on NVIDIA's VideoFDB perception track.
Griffin, described as the first Human Interaction Model to pass a video Turing Test, has already reached #1 on NVIDIA's full-duplex AI video benchmark, with 44% of people mistaking it for a real person compared to ~3% for other systems.
The author asserts that Opus 5.5 is the first AI model to personally pass the Turing test, marking a significant breakthrough towards AGI.
This paper distinguishes between aligning AI with human preferences versus human behavior, showing that preference alignment can reduce human-likeness and establishing a Turing-test gap in current alignment methods.
Phoenix-4.5 is introduced as the fastest and most expressive real-time human rendering model, featuring enhanced facial animation and upper body movement, claimed to be the closest AI to passing the Turing test face to face.
A paper inspired a game called TuringDuel where humans and AI compete in a one-word Turing test, revealing that humans lead AI 47–38 in wins, with 'poop' being an undefeated word choice.
Jurgen Schmidhuber argues that true AGI requires mastery of the real world and self-improving hardware, dismissing current claims as premature and criticizing the Turing Test as inadequate.
The article reflects on how AI's advancements, such as solving math problems and running local models, were once considered science fiction, urging appreciation for its current achievements.
The author reflects on how AI milestones like passing the Turing test and GPT-3's capabilities have been quickly normalized, urging the community to appreciate the current moment of AI advancement towards AGI.
This paper presents a formal impossibility result showing that tabular foundation models cannot distinguish legal from rule-violating database states without access to operational rules, verified through the Operational Turing Test (OTT) and empirical experiments.
This paper introduces Turing-RL, a reinforcement learning approach that uses Turing test-based rewards to train language models to generate responses indistinguishable from human users in conversational and forum settings, outperforming baseline methods.
This speculative paper argues that the Turing Test is flawed and that both human and AI intelligence are deterministic processes, with true consciousness requiring persistent memory loops and internal dialogue, framing AI as an evolutionary phase shift from carbon to silicon.
This paper introduces RogueAI, a reverse Turing test implemented as an interactive webapp where human players interrogate two LLM agents to identify which one is licensed to deceive within a shared fictional scenario. A pilot deployment shows a gap between heuristic detection (75.6% accuracy) and human performance (56.6%), highlighting the potential of the system as a data-collection and teaching tool for AI deception and honesty.
A research paper shows that while AI can solve CAPTCHAs as well as humans, behavioral differences in interaction patterns can still reliably distinguish bots from people, leading to the proposal of a 'Process Turing Test'.
A new study published in PNAS shows that advanced LLMs like GPT-4.5 can pass the Turing Test, with participants finding them more human than actual humans, prompting a reevaluation of what the test measures.
ClankerPass is a product that challenges users to convince an AI that they themselves are an AI, a reverse Turing test-style interaction.
Andrew Ng proposes a new "Turing-AGI Test" to better measure artificial general intelligence by having systems perform real work tasks with internet access, arguing that the term AGI has become overhyped and needs precise definition to avoid misleading stakeholders about AI capabilities.