Tag
This paper investigates the instability of large language model persona-driven generations in multiple-choice question answering (MCQA) tasks, proposing three metrics to measure performance, outcome, and correctness stability across model families, sizes, and question domains. The study finds that instability varies consistently, with math and commonsense questions showing greater instability, and that task prompt format introduces more instability than other hyperparameters like temperature.
This paper introduces Code-Guided Reasoning (CGR), an evaluation protocol for measuring how executable reasoning scaffolds improve small language model performance on multiple-choice question answering tasks, showing a significant accuracy improvement over direct answering.