Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons

Published in International Conference on Machine Learning, 2025

Using MathJax in the description is supported - \(E=mc^2\) - however, the use must be mindful that the default delimiters are $$...$$ and \\[...\\] which differs from the $...$ that is typically expected.

Recommended citation: Jianhui Chen, Yuzhang Luo, Liangming Pan. (2026). "Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units." ICML.
Download Paper