Tag
The author tests AI image detectors on compressed files and discusses how to interpret detection scores in real-world scenarios, highlighting conflicts between scores and visual inspection.
Evaluation of gpt-4o-mini and gpt-4o on an event classification system showed gpt-4o performed better, but both models had unreliable confidence scores for real-world decision-making.
The author argues that asking LLMs to produce self-assessed confidence scores is scientifically invalid and misleading, serving as a psychological safety trick rather than a reliable measure of correctness.
Vik Paruchuri announces research-driven safeguards that reduce OCR hallucinations to near-zero in their benchmark, with word-level bounding boxes and confidence scores for any remaining errors.