latent-space-probes

Tag

Cards List
#latent-space-probes

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

arXiv cs.LG · yesterday Cached

A reproducibility study of Khatri et al.'s latent-space safety probes, testing generalization across model families and sensitivity to non-determinism. Results show the probes extend to other models with similar F1 scores, and final token latent vectors remain consistent across seeds.

0 favorites 0 likes
← Back to home

Submit Feedback