Tag
The paper introduces a method using corpus characterization and inverse constitutional fine-tuning to improve the stylistic alignment of AI-generated radiology reports with authentic radiologist writing. This approach achieves significant gains in text alignment metrics, demonstrating effectiveness for style-aware report generation.
SynthSentry introduces a corpus-level, model-agnostic method to detect synthetic data contamination in language model training data without access to generating models, using distributional divergence over lexical, n-gram, and perplexity statistics.
SGHA is a fully automated system that uses a local 9B language model to discover research problems from scientific literature by structuring evidence and detecting structural gaps, offering transparency and privacy over proprietary models.
This paper analyzes whether LLM-generated novels exhibit compressed formal variation compared to human-written novels, finding that AI outputs are more uniform in sentence structure, readability, and punctuation despite individual novels resembling human style.
This paper presents a corpus-centric diagnostic framework for analyzing biomedical NER and EL benchmarks, revealing substantial differences across nine corpora and arguing that standard statistics are insufficient for characterizing evaluation demands.