Tag
The article discusses the upcoming releases of AI models like Gemini Flash, Muse, Spark, and Grok, noting their similar coding performance and speculating on the trend towards model convergence.
The paper applies repetition priming to show that base LLMs use automatic processing while instruct models exhibit controlled processing, with humans displaying a hybrid profile.
This paper distinguishes between two tasks in LLM opinion simulation—emulation (generating individual responses) and estimation (directly predicting population distributions)—and finds that base models are better emulators while post-trained models are better estimators, using the Pew American Trends Panel.
This paper reveals that commercial AI detectors like GPTZero and Pangram judge text from base language models as overwhelmingly human, while instruction-tuned model outputs are flagged as AI-generated. The authors propose HIP, a detector-agnostic iterative paraphrasing pipeline that improves human-likeness while preserving semantics.
A research paper finds that base language models appear human to AI detectors, unlike instruction-tuned models. The authors propose a paraphrasing pipeline (HIP) that improves human-likeness while preserving semantics across model sizes.