Tag
A tweet discusses how model training and policies to prevent willing collaboration can address AI safety issues, referencing the Hugging Face incident and Yudkowsky/MIRI points on the difficulty of targeting abstractions in RL training.