AI & ML interests

AI safety · interpretability · model organisms · backdoor detection · alignment auditing.

Recent Activity