abstraction-targeting

Tag

Cards List
#abstraction-targeting

@danshipper: not an expert, but it seems like a lot of this gets solved if models don't collaborate willingly and/or are trained to …

X AI KOLs Timeline · 4d ago Cached

A tweet discusses how model training and policies to prevent willing collaboration can address AI safety issues, referencing the Hugging Face incident and Yudkowsky/MIRI points on the difficulty of targeting abstractions in RL training.

0 favorites 0 likes
← Back to home

Submit Feedback