Check our @rajmovva.bsky.social 's earlier thread for details! One difference from our previous draft: we leave open the possibility that SAEs can also be a useful tool for acting on knowns (there is new evidence for this) bsky.app/profile/rajm...
Raj Movva📢New POSITION PAPER: Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts Despite recent results, SAEs aren't dead! They can still be useful to mech interp, and also much more broadly: across FAccT, computational social science, and ML4H. 🧵