Tag
A high Pearson correlation of 0.91 is found between a model's AA Intelligence Index score and its ability to generate Base64 encoded responses, despite no explicit training for that task.
This paper develops a minimal theory for partially correlated verifier cascades in LLM harnesses, showing concave log-odds, polynomial reliability decay, blind-spot ceilings, and that decorrelation is more effective than adding more gates.