Tag
This paper introduces STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets for psycholinguistic plausibility judgments. Experiments show that adding a global reasoning scratchpad and evaluator-guided refinement substantially improves generation quality, though near-boundary events remain challenging.
This position paper argues that LLM self-explanations can be plausible, questionably faithful, but highly actionable, and proposes evaluation guidelines beyond traditional metrics.
A tweet arguing that China's advantages in talent, scale, urgency, and fewer regulations make winning the AI race more plausible than commonly assumed.
This paper investigates the trade-off between plausibility and faithfulness in cross-lingual explanations from LLMs, finding that English-pivot explanations achieve higher span agreement with human rationales but suffer reduced causal faithfulness compared to native-language explanations.