Tag
This paper investigates how prompt robustness varies between objective and subjective questions in LLM evaluations, finding that sensitivity to prompt changes depends on question type, prompt change, and model.
This paper empirically demonstrates that single-prompt evaluation of instruction-tuned embedding models is insufficient, as performance varies significantly with prompt phrasing and leaderboard rankings can be manipulated by prompt selection.