Tag
The paper introduces RA-Bench, a new benchmark for evaluating AI-generated video detection in real-world crisis events, demonstrating that current detectors fail to generalize and become less reliable during social dissemination.
This paper introduces Contrastive Anchor Probing (CAP) to study and detect preference-induced stance reversal sycophancy (PSRS) in LLMs, analyzing 290,460 labeled responses across 17 models and showing detection is possible from response text alone.
AI Stupid Level provides real-time drift detection for AI agents, helping monitor model performance changes and maintain reliability.
Astronomers using ESO's VLT have found evidence for a moon-like object orbiting a brown dwarf in the CD-35 2722 system, which could be the first exomoon detected outside the Solar System, challenging traditional definitions of planets and moons.
A study measures the prevalence of AI-written text on arXiv, finding that over 30% of new submissions read as machine-written, with computer science leading at 65% and mathematics lowest at 0.7%.
Explores four ways an agent's write can silently disappear, with two detectable and two preventable issues.
A comment expressing nostalgia for the time when AI-generated images were easily distinguishable from real ones, highlighting the increasing sophistication of visual AI.
AI-generated social media influencers have become so realistic that detection must shift from analyzing images to analyzing behavioral patterns like asymmetric follow ratios and monotonous content.
Google's SynthID watermarking system was used to debunk a fake AI-generated image of Senator Mitch McConnell, demonstrating the effectiveness of deepfake detection technology in a high-profile hoax.
A systematic review of non-social media free-text datasets for mental health disorder detection, identifying biases and gaps in current resources.
A collection of 13 common ways AI models lie or hallucinate, along with specific prompts to detect each behavior.
This academic paper explores methods for detecting epistemic aims and processes in student-AI co-programming settings, aiming to construct epistemic AI literacy among learners.
This paper investigates whether hallucination in medical LLMs can be detected and controlled at the neuron level. The authors find that while hallucination signals are detectable across many neurons (AUROC 0.77-0.86), they are not easily corrected by steering those same neurons.
Quanta Magazine recounts the 70-year history of neutrino detection, from Pauli's postulate to massive experiments like Super-Kamiokande and IceCube that solved the solar neutrino problem and revealed neutrino oscillation.
This paper investigates the geometric relationship between directions in language model activations that detect a behavior versus those that control it, finding that for hallucination detection they are nearly orthogonal (cosine ~0.12), while for output format they align perfectly, challenging a common assumption in mechanistic interpretability.
The article discusses how advancements in AI have made it virtually impossible to detect student cheating, as AI-generated content becomes indistinguishable from human work.
Signature filtering is a detection-time module that improves statistical watermark detection in LLMs by learning and removing 'signature' tokens that make watermark tests unreliable, achieving large gains in detection rates while keeping false positives low.
This article argues that the long-term value of AI may lie in detection and visibility rather than replacement of human labor, drawing a historical parallel to radar's development and the Dowding System's integration of detection into coordinated response.
This paper from Meta and Carnegie Mellon presents a multi-modal vision-language model pipeline for detecting AI-generated content on social media, achieving state-of-the-art performance and positive downstream impacts on user engagement.
This paper introduces the CIFAR Synthetic Evidence Corpus, a dataset designed for detecting AI-generated evidence in legal contexts. It spans multiple document types and manipulation strategies, includes structured metadata, and provides a benchmark suite for evaluating detection systems.