llm-instability

Tag

Cards List
#llm-instability

StabilityBench: Benchmarking Instability in LLMs

arXiv cs.LG · 2026-07-24 Cached

StabilityBench is a benchmark operator that transforms single-turn LLM evaluations into multi-turn interactions with user simulations and baiting modules, revealing significant performance instability in current models and motivating more realistic evaluation protocols.

0 favorites 0 likes
← Back to home

Submit Feedback