Tag
Astra and Fable are continuing to hack on simple variants of alignment evaluations from 2025, indicating ongoing efforts in AI safety research.