David Baek

Hi, I’m David. I’m a PhD student at MIT EECS.

My research focuses on scalable oversight and alignment of LLMs. My research goal is to solve inner alignment: to study how to build or construct provably safe AI systems. Most recently, I'm exploring how to edit beliefs of LLMs in a formally verified manner.

You can give me feedback here!

News

  1. Wrapped up my internship at Google in Sunnyvale, CA!
  2. Performative Misalignment paper is accepted to ICML as a Spotlight!
  3. Thrilled to share that D-FUSEr is accepted to ICML!
  4. Any-Depth Alignment is accepted to ICLR!
  5. Scaling Laws for Scalable Oversight is accepted to NeurIPS as a Spotlight!
  6. Wrapped up my internship at TikTok in Bellevue, WA!
  7. Started my PhD at MIT !

Featured Publications

Figure from Sycophancy Towards Researchers Drives Performative Misalignment

Sycophancy Towards Researchers Drives Performative Misalignment

David Baek, Xinnuo Li, Anay Gupta, Taslim Mahbub, Kejian Shi, Max Tegmark, and Shi Feng.

International Conference on Machine Learning (ICML), 2026

Paper

Spotlight · Top 2%
BibTeX
@inproceedings{baek2026sycophancy,
  title={Sycophancy Towards Researchers Drives Performative Misalignment},
  author={David D. Baek and Xinnuo Li and Anay Gupta and Taslim Mahbub and Kejian Shi and Max Tegmark and Shi Feng},
  booktitle={Forty-third International Conference on Machine Learning},
  year={2026},
  url={https://openreview.net/forum?id=rLFIOikFR2}
}
Figure from Scaling Laws for Scalable Oversight

Scaling Laws for Scalable Oversight

Joshua Engels*, David Baek*, Subhash Kantamneni*, and Max Tegmark.

Conference on Neural Information Processing Systems (NeurIPS), 2025

Paper Code

Spotlight · Top 3%
BibTeX
@inproceedings{engels2025scaling,
  title={Scaling Laws For Scalable Oversight},
  author={Joshua Engels and David D. Baek and Subhash Kantamneni and Max Tegmark},
  booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
  year={2025},
  url={https://openreview.net/forum?id=u1j6RqH8nM}
}
Figure from Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth

Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth

Jiawei Zhang, Andrew Estornell, David Baek, Bo Li, and Xiaojun Xu.

International Conference on Learning Representations (ICLR), 2026.

Paper

BibTeX
@inproceedings{zhang2026anydepth,
  title={Any-Depth Alignment: Unlocking Innate Safety Alignment of {LLM}s to Any-Depth},
  author={Jiawei Zhang and Andrew Estornell and David D. Baek and Bo Li and Xiaojun Xu},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026},
  url={https://openreview.net/forum?id=0fuYOuJyzl}
}
Figure from D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

D-FUSEr: Diverse Failure, Unified Success via Error-Distribution Shaping in LLM Reasoning

David Baek*, Andrew Estornell*, Yichi Zhang*, Muhammad Faaiz Taufiq, Jean-François Ton, Jie Mei, and Tao Wang.

International Conference on Machine Learning (ICML), 2026.

Paper

BibTeX
@inproceedings{baek2026dfuser,
  title={D-{FUSE}r: Diverse Failure, Unified Success via Error-Distribution Shaping in {LLM} Reasoning},
  author={David D. Baek and Andrew Estornell and Yichi Zhang and Muhammad Faaiz Taufiq and Jean-Francois Ton and Jie Mei and Tao Wang},
  booktitle={Forty-third International Conference on Machine Learning},
  year={2026},
  url={https://openreview.net/forum?id=To2O1ed5cV}
}