John Kirchenbauer

portrait.png

I’m a postdoctoral fellow at the Vector Institute, working with Colin Raffel on building open language models and understanding and mitigating the risks posed by modern AI. I earned my PhD at the University of Maryland advised by Tom Goldstein. Previously, I interned at Google X and worked at Carnegie Mellon’s Software Engineering Institute. I also studied computer science at Washington University in St. Louis and violin at Oberlin Conservatory.

Research directions

  • Efficient learning and inference. Distillation, multi-token prediction, and methods for making language models more efficient.
  • Security and provenance. Watermarks and fingerprints for identifying generated content and protecting models and datasets.
  • Memorization and privacy. Understanding what language models learn, retain, and reproduce from their training data.
  • Open language models and data. Openly licensed training corpora, model development, and the relationship between data and model behavior.

Browse all publications and preprints →

News

Aug 04, 2026 Started my Postdoc at the Vector Institute in Toronto with Colin Raffel.
May 06, 2026 Defended my Dissertation 🎉: “Enhancing Trust and Transparency in Language Model Development and Deployment”.

Selected publications

A Watermark for Large Language Models

John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein

ICML 2023 · 2023
Outstanding Paper Award, ICML 2023

A method for marking generated text so that its origin can be detected statistically.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

Nikhil Kandpal, Brian Lester, Colin Raffel, Sebastian Majstorovic, Stella Biderman, Baber Abbasi , … John Kirchenbauer
All 27 authors

Nikhil Kandpal, Brian Lester, Colin Raffel, Sebastian Majstorovic, Stella Biderman, Baber Abbasi, Luca Soldaini, Enrico Shippole, A. Feder Cooper, Aviya Skowron, John Kirchenbauer, Shayne Longpre, Lintang Sutawika, Alon Albalak, Zhenlin Xu, Guilherme Penedo, Loubna Ben Allal, Elie Bakouch, John David Pressman, Honglu Fan, Dashiell Stander, Guangyu Song, Aaron Gokaslan, Tom Goldstein, Brian R. Bartoldson, Bhavya Kailkhura, Tyler Murray

arXiv preprint · 2025

An open training corpus built from public-domain and openly licensed text.

Multi-Token Prediction via Self-Distillation

John Kirchenbauer, Abhimanyu Hans, Brian Bartoldson, Micah Goldblum, Ashwinee Panda, Tom Goldstein

arXiv preprint · 2026

Training language models to predict multiple future tokens using self-distillation.