Publications & Presentations
Key research outputs. First-author publications are highlighted.
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
Accepted to Findings of EMNLP 2026.
The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
Spotlight at Mechanistic Interpretability Workshop, NeurIPS 2025.
Self-Ablating Transformers: More Interpretability, Less Sparsity
Poster at Building Trust Workshop, ICLR 2025.
World Model Agents with Change-Based Intrinsic Motivation
Oral presentation at NLDL 2025.