@panickssery: In addition to identifying LLM use, Pangram has identified many academics who don't know the difference between false p…
Summary
Pangram, a tool for detecting LLM use, reveals that many academics do not understand the difference between false positives and false negatives.
Similar Articles
@rohanpaul_ai: LLMs can accept the same false claim differently depending on its tone, certainty, and grammatical form. Small wording …
This paper presents EoBench, a benchmark testing LLMs' acceptance of false claims phrased in 19 different styles, finding that tone, certainty, and grammatical form significantly affect model responses, with larger and instruction-tuned models showing more resistance.
Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content
This paper investigates how well LLMs can detect their own generated content in educational contexts, finding that detection accuracy varies by task type and is unreliable for short-answer questions.
@rohanpaul_ai: LLMs often cannot tell when an attack made them say something unsafe. Asking an LLM whether its own previous answer was…
This paper investigates whether LLMs can reliably self-report when their outputs have been compromised by adversarial prefills, finding that models often cannot distinguish between compromised and intentional outputs, and their limited recognition stems from normal refusal behavior rather than true self-awareness.
It's Not Just X. It's Y
This essay explores how LLMs' post-training (RLHF and RLVR) produces linguistic tics like negative parallelism, and critiques the use of AI detectors (Grammarly, Pangram) that force writers to sound like machines to avoid false accusations.
@omarsar0: LLM review weirdness indeed. Avoid using scores with LLM judges, or be extremely careful if you do. Use binary labels w…
A tweet discussing a discovered quirk where renaming a paper PDF to a longer, positive title improves LLM judge scores, advising caution with score-based LLM evaluation and recommending binary labels instead.