John Kirchenbauer
I’m a postdoctoral fellow at the Vector Institute, working with Colin Raffel on building open language models and understanding and mitigating the risks posed by modern AI. I earned my PhD at the University of Maryland advised by Tom Goldstein. Previously, I interned at Google X and worked at Carnegie Mellon’s Software Engineering Institute. I also studied computer science at Washington University in St. Louis and violin at Oberlin Conservatory.
Research directions
- Efficient learning and inference. Distillation, multi-token prediction, and methods for making language models more efficient.
- Security and provenance. Watermarks and fingerprints for identifying generated content and protecting models and datasets.
- Memorization and privacy. Understanding what language models learn, retain, and reproduce from their training data.
- Open language models and data. Openly licensed training corpora, model development, and the relationship between data and model behavior.
News
| Aug 04, 2026 | Started my Postdoc at the Vector Institute in Toronto with Colin Raffel. |
|---|---|
| May 06, 2026 | Defended my Dissertation 🎉: “Enhancing Trust and Transparency in Language Model Development and Deployment”. |
Selected publications
A Watermark for Large Language Models
ICML 2023 · 2023
Outstanding Paper Award, ICML 2023
A method for marking generated text so that its origin can be detected statistically.
Abstract
Potential harms of large language models can be mitigated by watermarking model output, i.e., embedding signals into generated text that are invisible to humans but algorithmically detectable from a short span of tokens. We propose a watermarking framework for proprietary language models. The watermark can be embedded with negligible impact on text quality, and can be detected using an efficient open-source algorithm without access to the language model API or parameters. The watermark works by selecting a randomized set of "green" tokens before a word is generated, and then softly promoting use of green tokens during sampling. We propose a statistical test for detecting the watermark with interpretable p-values, and derive an information-theoretic framework for analyzing the sensitivity of the watermark. We test the watermark using a multi-billion parameter model from the Open Pretrained Transformer (OPT) family, and discuss robustness and security.
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
arXiv preprint · 2025
An open training corpus built from public-domain and openly licensed text.
Abstract
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but prior data collection efforts have yielded datasets too small or low-quality to produce performant LLMs. To address this gap, we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining. The Common Pile comprises content from 30 sources that span diverse domains including research papers, code, books, encyclopedias, educational materials, audio transcripts, and more. Crucially, we validate our efforts by training two 7 billion parameter LLMs on text from the Common Pile: Comma v0.1-1T and Comma v0.1-2T, trained on 1 and 2 trillion tokens respectively. Both models attain competitive performance to LLMs trained on unlicensed text with similar computational budgets, such as Llama 1 and 2 7B. In addition to releasing the Common Pile v0.1 itself, we also release the code used in its creation as well as the training mixture and checkpoints for the Comma v0.1 models.
Multi-Token Prediction via Self-Distillation
arXiv preprint · 2026
Training language models to predict multiple future tokens using self-distillation.
Abstract
Existing techniques for accelerating language model inference, such as speculative decoding, require training auxiliary speculator models and building and deploying complex inference pipelines. We consider a new approach for converting a pretrained autoregressive language model from a slow single next token prediction model into a fast standalone multi-token prediction model using a simple online distillation objective. The final model retains the exact same implementation as the pretrained initial checkpoint and is deployable without the addition of any auxiliary verifier or other specialized inference code. Our method produces models that decode more than $3\times$ faster at $<5\%$ drop in accuracy on GSM8K relative to the single token decoding performance of the same checkpoint.