Tag
This paper introduces BenchDrift, a method for quantifying how LLM benchmark performance changes when problems are rephrased without changing meaning or answer. It shows that rephrasing causes bidirectional correctness flips across models and benchmarks, with stronger models becoming more sensitive to phrasing.
Describes PHI // DRIFT, a cognitive architecture with seven homeostatic state variables that drift between sessions, memory scored by emotional salience and time decay, and a Jungian shadow module, built on a CPU-only mini tower and submitted as a preprint to SSRN.