mcqa

Tag

Cards List
#mcqa

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

arXiv cs.CL · 2026-07-02 Cached

This paper investigates the instability of large language model persona-driven generations in multiple-choice question answering (MCQA) tasks, proposing three metrics to measure performance, outcome, and correctness stability across model families, sizes, and question domains. The study finds that instability varies consistently, with math and commonsense questions showing greater instability, and that task prompt format introduces more instability than other hyperparameters like temperature.

0 favorites 0 likes
#mcqa

Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffolds

Hugging Face Daily Papers · 2026-05-12 Cached

This paper introduces Code-Guided Reasoning (CGR), an evaluation protocol for measuring how executable reasoning scaffolds improve small language model performance on multiple-choice question answering tasks, showing a significant accuracy improvement over direct answering.

0 favorites 0 likes
← Back to home

Submit Feedback