Publications & Presentations

Key research outputs. First-author publications are highlighted.

Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes

Jeremias Lino Ferrao, Niclas Müller-Hof, Iustin Sîrbu, Traian Rebedea, and Yftah Ziser

The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features

Jeremias Lino Ferrao, Matthijs van der Lende, Ilija Lichkovski, and Clement Neo

Self-Ablating Transformers: More Interpretability, Less Sparsity

Jeremias Lino Ferrao, Luhan Mikaelson, Keenan Pepper, and Natalia Perez-Campanero Antolin

World Model Agents with Change-Based Intrinsic Motivation

Jeremias Lino Ferrao and Rafael F. Cunha