Tag
A comment expressing nostalgia for the time when AI-generated images were easily distinguishable from real ones, highlighting the increasing sophistication of visual AI.
AI-generated social media influencers have become so realistic that detection must shift from analyzing images to analyzing behavioral patterns like asymmetric follow ratios and monotonous content.
Google's SynthID watermarking system was used to debunk a fake AI-generated image of Senator Mitch McConnell, demonstrating the effectiveness of deepfake detection technology in a high-profile hoax.
A systematic review of non-social media free-text datasets for mental health disorder detection, identifying biases and gaps in current resources.
A collection of 13 common ways AI models lie or hallucinate, along with specific prompts to detect each behavior.
This academic paper explores methods for detecting epistemic aims and processes in student-AI co-programming settings, aiming to construct epistemic AI literacy among learners.
This paper investigates whether hallucination in medical LLMs can be detected and controlled at the neuron level. The authors find that while hallucination signals are detectable across many neurons (AUROC 0.77-0.86), they are not easily corrected by steering those same neurons.
Quanta Magazine recounts the 70-year history of neutrino detection, from Pauli's postulate to massive experiments like Super-Kamiokande and IceCube that solved the solar neutrino problem and revealed neutrino oscillation.
This paper investigates the geometric relationship between directions in language model activations that detect a behavior versus those that control it, finding that for hallucination detection they are nearly orthogonal (cosine ~0.12), while for output format they align perfectly, challenging a common assumption in mechanistic interpretability.
The article discusses how advancements in AI have made it virtually impossible to detect student cheating, as AI-generated content becomes indistinguishable from human work.
Signature filtering is a detection-time module that improves statistical watermark detection in LLMs by learning and removing 'signature' tokens that make watermark tests unreliable, achieving large gains in detection rates while keeping false positives low.
This article argues that the long-term value of AI may lie in detection and visibility rather than replacement of human labor, drawing a historical parallel to radar's development and the Dowding System's integration of detection into coordinated response.
This paper from Meta and Carnegie Mellon presents a multi-modal vision-language model pipeline for detecting AI-generated content on social media, achieving state-of-the-art performance and positive downstream impacts on user engagement.
This paper introduces the CIFAR Synthetic Evidence Corpus, a dataset designed for detecting AI-generated evidence in legal contexts. It spans multiple document types and manipulation strategies, includes structured metadata, and provides a benchmark suite for evaluating detection systems.
User expresses confusion over distinguishing real from AI-generated video, highlighting the growing realism of synthetic media.
A large-scale empirical study analyzes 284 linguistic features across 27 LLMs and 10 text domains to assess which features reliably detect AI-generated text. The study finds that lexical richness measures are the most robust cross-domain and cross-model signals, while many other proposed indicators are strongly context-dependent.
This paper investigates whether open-source quantized LLMs encode a linearly separable truthfulness signal in their hidden states. Across three 7B-8B instruction-tuned models, a linear probe on a single mid-network layer achieves 0.904-1.000 AUROC on hallucination detection benchmarks, outperforming sampling-based methods.
The article explores the current state of AI-generated writing, its detection, and the implications for education and literature, referencing the Granta controversy where a story suspected to be AI-written won a prize.
Introduces LUNA, a linguistics-aware LLM watermarking method that achieves non-distortionary embedding and model-free detection across multiple languages, significantly improving AUROC and perplexity preservation.
Introduces SynCred-Bench, a benchmark of 600 AI-generated misinformation images across six credible-form categories, showing that existing detectors (including MLLMs, open-source AIGC detectors, and commercial APIs) perform poorly, with human annotators also struggling.