About Me

I am an incoming Ph.D. student in Computer Science. I graduated from the Knowledge Engineering Group (KEG) at Tsinghua University in 2025, where I was fortunate to be advised by Prof. Juanzi Li, Prof. Lei Hou and Prof. Xiaozhi Wang. Currently, I am a research assistant at the PILLAR Lab at Peking University, working with Prof. Liangming Pan on understanding the mechanisms of large language models.

My long-term goal is to develop a scientific understanding of intelligent systems and leverage such understanding to build more capable, trustworthy, and interpretable AI. My current research interests include:

  • Mechanistic Interpretability
    • Reverse-engineering neural networks to identify interpretable computational structures. Understanding how complex capabilities emerge from model parameters and representations.
  • Large Language Model Reasoning
    • Studying the internal dynamics and representations of reasoning processes. Exploring the computational foundations of reasoning abilities in LLMs.
  • Data Attribution
    • Understanding how training data and optimization shape model behaviors.

News

2026/05
Paper Our paper Mechanistic Data Attribution was accepted by ICML 2026 and selected as Oral Presentation (Top 0.7%).
2026/04
Paper Our survey paper on Large Reasoning Models Mechanism was accepted by ACL 2026.
2025/09
Paper Our paper Safety Neurons was accepted by NeurIPS 2025.

Selected Publications

2026
ICML 2026 Oral

Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

Jianhui Chen*, Yuzhang Luo*, Liangming Pan

2026
ACL 2026 Oral

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

Yi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao, Hao Peng, Xiaobao Wu, Jianhui Chen, Muhan Zhang, Liangming Pan

2025
NeurIPS 2025

Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons

Jianhui Chen*, Xiaozhi Wang*, Zijun Yao, Yushi Bai, Lei Hou, Juanzi Li

Academic Service

Reviewer

  • ACL 2026, ICML 2026, NeurIPS 2026