Publications

See my Google Scholar profile for the most up-to-date list and citation information.

2026

  1. Watermarking for Proprietary Dataset Protection

    John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Tom Goldstein

    arXiv preprint · 2026

  2. Multi-Token Prediction via Self-Distillation

    John Kirchenbauer, Abhimanyu Hans, Brian Bartoldson, Micah Goldblum, Ashwinee Panda, Tom Goldstein

    arXiv preprint · 2026

  3. Antidistillation Fingerprinting

    Yixuan Even Xu, John Kirchenbauer, Yash Savani, Asher Trockman, Alexander Robey, Tom Goldstein
    All 8 authors

    Yixuan Even Xu, John Kirchenbauer, Yash Savani, Asher Trockman, Alexander Robey, Tom Goldstein, Fei Fang, J. Zico Kolter

    arXiv preprint · 2026

2025

  1. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence

    Sean McLeish, Ang Li, John Kirchenbauer, Dayal Singh Kalra, Brian R. Bartoldson, Bhavya Kailkhura
    All 10 authors

    Sean McLeish, Ang Li, John Kirchenbauer, Dayal Singh Kalra, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Jonas Geiping, Tom Goldstein, Micah Goldblum

    arXiv preprint · 2025

  2. The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text

    Nikhil Kandpal, Brian Lester, Colin Raffel, Sebastian Majstorovic, Stella Biderman, Baber Abbasi , … John Kirchenbauer
    All 27 authors

    Nikhil Kandpal, Brian Lester, Colin Raffel, Sebastian Majstorovic, Stella Biderman, Baber Abbasi, Luca Soldaini, Enrico Shippole, A. Feder Cooper, Aviya Skowron, John Kirchenbauer, Shayne Longpre, Lintang Sutawika, Alon Albalak, Zhenlin Xu, Guilherme Penedo, Loubna Ben Allal, Elie Bakouch, John David Pressman, Honglu Fan, Dashiell Stander, Guangyu Song, Aaron Gokaslan, Tom Goldstein, Brian R. Bartoldson, Bhavya Kailkhura, Tyler Murray

    arXiv preprint · 2025

  3. FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition

    John Kirchenbauer, Janny Mongkolsupawan, Yuxin Wen, Tom Goldstein, Daphne Ippolito

    arXiv preprint · 2025

  4. Zero-Shot Vision Encoder Grafting via LLM Surrogates

    Kaiyu Yue, Vasu Singla, Menglin Jia, John Kirchenbauer, Rifaa Qadri, Zikui Cai
    All 9 authors

    Kaiyu Yue, Vasu Singla, Menglin Jia, John Kirchenbauer, Rifaa Qadri, Zikui Cai, Abhinav Bhatele, Furong Huang, Tom Goldstein

    arXiv preprint · 2025

  5. When Can You Get Away with Low Memory Adam?

    Dayal Singh Kalra, John Kirchenbauer, Maissam Barkeshli, Tom Goldstein

    arXiv preprint · 2025

  6. Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers

    Siddharth Singh, Prajwal Singhania, Aditya Ranjan, John Kirchenbauer, Jonas Geiping, Yuxin Wen
    All 12 authors

    Siddharth Singh, Prajwal Singhania, Aditya Ranjan, John Kirchenbauer, Jonas Geiping, Yuxin Wen, Neel Jain, Abhimanyu Hans, Manli Shu, Aditya Tomar, Tom Goldstein, Abhinav Bhatele

    arXiv preprint · 2025

  7. Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs

    Ryan Synk, Monte Hoover, John Kirchenbauer, Neel Jain, Alex Stein, Manli Shu
    All 9 authors

    Ryan Synk, Monte Hoover, John Kirchenbauer, Neel Jain, Alex Stein, Manli Shu, Josue Melendez Sanchez, Ramani Duraiswami, Tom Goldstein

    arXiv preprint · 2025

  8. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

    Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson
    All 9 authors

    Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein

    arXiv preprint · 2025

  9. Gemstones: A Model Suite for Multi-Faceted Scaling Laws

    Sean McLeish, John Kirchenbauer, David Yu Miller, Siddharth Singh, Abhinav Bhatele, Micah Goldblum
    All 8 authors

    Sean McLeish, John Kirchenbauer, David Yu Miller, Siddharth Singh, Abhinav Bhatele, Micah Goldblum, Ashwinee Panda, Tom Goldstein

    arXiv preprint · 2025

2024

  1. Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs

    Abhimanyu Hans, Yuxin Wen, Neel Jain, John Kirchenbauer, Hamid Kazemi, Prajwal Singhania
    All 11 authors

    Abhimanyu Hans, Yuxin Wen, Neel Jain, John Kirchenbauer, Hamid Kazemi, Prajwal Singhania, Siddharth Singh, Gowthami Somepalli, Jonas Geiping, Abhinav Bhatele, Tom Goldstein

    arXiv preprint · 2024

  2. GenQA: Generating Millions of Instructions from a Handful of Prompts

    Jiuhai Chen, Rifaa Qadri, Yuxin Wen, Neel Jain, John Kirchenbauer, Tianyi Zhou
    All 7 authors

    Jiuhai Chen, Rifaa Qadri, Yuxin Wen, Neel Jain, John Kirchenbauer, Tianyi Zhou, Tom Goldstein

    arXiv preprint · 2024

  3. OPTune: Efficient Online Preference Tuning

    Lichang Chen, Jiuhai Chen, Chenxi Liu, John Kirchenbauer, Davit Soselia, Chen Zhu
    All 9 authors

    Lichang Chen, Jiuhai Chen, Chenxi Liu, John Kirchenbauer, Davit Soselia, Chen Zhu, Tom Goldstein, Tianyi Zhou, Heng Huang

    arXiv preprint · 2024

  4. Transformers Can Do Arithmetic with the Right Embeddings

    Sean McLeish, Arpit Bansal, Alex Stein, Neel Jain, John Kirchenbauer, Brian R. Bartoldson
    All 11 authors

    Sean McLeish, Arpit Bansal, Alex Stein, Neel Jain, John Kirchenbauer, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Jonas Geiping, Avi Schwarzschild, Tom Goldstein

    arXiv preprint · 2024

  5. LMD3: Language Model Data Density Dependence

    John Kirchenbauer, Garrett Honke, Gowthami Somepalli, Jonas Geiping, Daphne Ippolito, Katherine Lee
    All 8 authors

    John Kirchenbauer, Garrett Honke, Gowthami Somepalli, Jonas Geiping, Daphne Ippolito, Katherine Lee, Tom Goldstein, David Andre

    arXiv preprint · 2024

2023

  1. NEFTune: Noisy Embeddings Improve Instruction Finetuning

    Neel Jain, Ping-yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli
    All 13 authors

    Neel Jain, Ping-yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, Micah Goldblum, Jonas Geiping, Tom Goldstein

    arXiv preprint · 2023

  2. Baseline Defenses for Adversarial Attacks Against Aligned Language Models

    Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang
    All 10 authors

    Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, Tom Goldstein

    arXiv preprint · 2023

  3. Bring Your Own Data! Self-Supervised Evaluation for Large Language Models

    Neel Jain, Khalid Saifullah, Yuxin Wen, John Kirchenbauer, Manli Shu, Aniruddha Saha
    All 9 authors

    Neel Jain, Khalid Saifullah, Yuxin Wen, John Kirchenbauer, Manli Shu, Aniruddha Saha, Micah Goldblum, Jonas Geiping, Tom Goldstein

    arXiv preprint · 2023

  4. On the Reliability of Watermarks for Large Language Models

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong
    All 10 authors

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, Tom Goldstein

    arXiv preprint · 2023

  5. Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust

    Yuxin Wen, John Kirchenbauer, Jonas Geiping, Tom Goldstein

    arXiv preprint · 2023

  6. Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery

    Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, Tom Goldstein

    arXiv preprint · 2023

  7. A Watermark for Large Language Models

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein

    ICML 2023 · 2023
    Outstanding Paper Award, ICML 2023

2022

  1. What is Your Metric Telling You? Evaluating Classifier Calibration under Context-Specific Definitions of Reliability

    John Kirchenbauer, Jacob Oaks, Eric Heim

    arXiv preprint · 2022