publications

L3 Lab publications.

2026

  1. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
    Anmol Agarwal, Natalie Neamtu, Pranjal Aggarwal, and 6 more authors
    arXiv, 2026
  2. ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
    Riyaz AhujaTate Rowney, Jeremy Avigad, and 1 more author
    arXiv, 2026
  3. On the Limits and Opportunities of AI Reviewers: Reviewing the Reviews of Nature-Family Papers with 45 Expert Scientists
    Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, and 55 more authors
    arXiv, 2026
  4. Reinforcing Human Behavior Simulation via Verbal Feedback
    Weiwei Sun, Xuhui Zhou, Jiarui Liu, and 13 more authors
    arXiv, 2026
  5. Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
    Guijin Son, Seungone Kim, Catherine Arnett, and 73 more authors
    arXiv, 2026
  6. AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
    Weihua Du, Jingming Zhuo, Yixin Dong, and 9 more authors
    arXiv, 2026
  7. Gym-Anything: Turn any Software into an Agent Environment
    Pranjal Aggarwal, Graham Neubig, and Sean Welleck
    arXiv, 2026
  8. Explorable Theorems: Making Written Theorems Explorable by Grounding Them in Formal Representations
    Hita Kambhamettu, Will Crichton, Sean Welleck, and 2 more authors
    arXiv, 2026
  9. Reasoning over Mathematical Objects: On-Policy Reward Modeling and Test Time Aggregation
    Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim, and 18 more authors
    arXiv, 2026
  10. Argument Reconstruction as Supervision for Critical Thinking in LLMs
    Hyun Ryu, Gyouk Chu, Gregor Betz, and 3 more authors
    arXiv, 2026
  11. Reasoning with Latent Tokens in Diffusion Language Models
    Andre HeSean Welleck, and Daniel Fried
    arXiv, 2026
  12. DSLean: A Framework for Type-Correct Interoperability Between Lean 4 and External DSLs
    Tate RowneyRiyaz Ahuja, Jeremy Avigad, and 1 more author
    2026
    FMCAD 2026
  13. GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
    Ningyuan Yang, Weihua Du, Weiwei Sun, and 2 more authors
    2026
    COLM 2026
  14. Mind the Sim2Real Gap in User Simulation for Agentic Tasks
    Xuhui Zhou, Weiwei Sun, Qianou Ma, and 8 more authors
    2026
    COLM 2026
  15. Training Proactive and Personalized LLM Agents
    Weiwei Sun, Xuhui Zhou, Weihua Du, and 5 more authors
    2026
    COLM 2026
  16. Propose, Solve, Verify: Self-Play Through Formal Verification
    Alex Wilf, Pranjal Aggarwal, Bryan Parno, and 4 more authors
    2026
    ICML 2026
  17. LeanArchitect: Automating Blueprint Generation for Humans and AI
    Thomas Zhu, Pietro Monticone, Jeremy Avigad, and 1 more author
    2026
    ITP 2026
  18. Scaling Evaluation-Time Compute with Reasoning Models as Evaluators
    Seungone Kim, Ian Wu, Jinu Lee, and 8 more authors
    In Findings of the Association for Computational Linguistics: ACL 2026, 2026
    ACL 2026 Findings
  19. Premise Selection for a Lean Hammer
    Thomas Zhu, Joshua Clune, Jeremy Avigad, and 2 more authors
    2026
    ICLR 2026 Oral (Top 1%)
  20. Programming with Pixels: Can Computer-Use Agents do Software Engineering?
    Pranjal Aggarwal, and Sean Welleck
    2026
    ICLR 2026
  21. RefineBench: Evaluating Refinement Capability of Language Models via Checklists
    Young-Jun Lee, Seungone Kim, Byung-Kwan Lee, and 6 more authors
    2026
    ICLR 2026
  22. OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
    Pranjal AggarwalSeungone Kim, Jack Lanchantin, and 4 more authors
    2026
    ICLR 2026
  23. The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
    Seongyun Lee, Seungone Kim, Minju Seo, and 9 more authors
    2026
    ICLR 2026

2025

  1. Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
    Andre He, Daniel Fried, and Sean Welleck
    2025
    EMNLP 2025 Oral
  2. Agentic-R1: Distilled Dual-Strategy Reasoning
    Weihua DuPranjal AggarwalSean Welleck, and 1 more author
    arXiv, 2025
    EMNLP 2025
  3. The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
    Seungone Kim, Juyoung Suk, Ji Yong Cho, and 29 more authors
    In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025
    Best Paper Award
  4. Evaluating Language Models as Synthetic Data Generators
    Seungone Kim, Juyoung Suk, Xiang Yue, and 7 more authors
    In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2025
  5. Optimizing Temperature for Language Models with Multi-Sample Inference
    Weihua Du, Yiming Yang, and Sean Welleck
    arXiv preprint arXiv:2502.05234, Jul 2025
    ICML 2025
  6. Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
    Yangzhen Wu, Zhiqing Sun, Shanda Li, and 2 more authors
    Jul 2025
    ICLR 2025
  7. miniCTX: Neural Theorem Proving with (Long-)Contexts
    Jiewen HuThomas Zhu, and Sean Welleck
    ArXiv, Jul 2025
    ICLR 2025 Oral
  8. ImProver: Agent-Based Automated Proof Optimization
    Riyaz Ahuja, Jeremy Avigad, Prasad Tetali, and 1 more author
    Jul 2025
    ICLR 2025
  9. Lean-STaR: Learning to Interleave Thinking and Proving
    Haohan Lin, Zhiqing Sun, Yiming Yang, and 1 more author
    Jul 2025
    ICLR 2025 Spotlight

2024

  1. Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
    Zhiqing Sun, Longhui Yu, Yikang Shen, and 4 more authors
    In Advances in Neural Information Processing Systems, Jul 2024
  2. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
    Seungone Kim, Juyoung Suk, Shayne Longpre, and 7 more authors
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024
  3. From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
    Sean Welleck, Amanda Bertsch, Matthew Finlayson, and 5 more authors
    Transactions on Machine Learning Research, Oct 2024
  4. miniCodeProps: a Minimal Benchmark for Proving Code Properties
    E. Lohn, and Sean Welleck
    Oct 2024
    NeurIPS 2024 Workshop on Safe Generative AI
  5. Llemma: An Open Language Model For Mathematics
    Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, and 6 more authors
    In International Conference on Learning Representations, Oct 2024
    ICLR 2024

2023

  1. A Survey of Deep Learning for Mathematical Reasoning
    Pan Lu, Liang Qiu, Wenhao Yu, and 2 more authors
    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2023
  2. MAUVE Scores for Generative Models: Theory and Practice
    Krishna Pillutla, Lang Liu, John Thickstun, and 6 more authors
    JMLR, Jul 2023
  3. Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning
    Ximing Lu, Faeze Brahman, Peter West, and 14 more authors
    EMNLP, Jul 2023
  4. STEER: Unified Style Transfer with Expert Reinforcement
    Skyler Hallinan, Faeze Brahman, Ximing Lu, and 3 more authors
    EMNLP Findings, Jul 2023
  5. Self-Refine: Iterative Refinement with Self-Feedback
    Aman Madaan, Niket Tandon, Prakhar Gupta, and 13 more authors
    In Thirty-seventh Conference on Neural Information Processing Systems, Jul 2023
  6. Generating Sequences by Learning to Self-Correct
    Sean Welleck, Ximing Lu, Peter West, and 4 more authors
    In The Eleventh International Conference on Learning Representations , Jul 2023
  7. llmstep: LLM proofstep suggestions in Lean
    Sean Welleck, and Rahul Saha
    In The 3rd Workshop on Mathematical Reasoning and AI at NeurIPS’23, Jul 2023
  8. Neural theorem proving tutorial
    Sean Welleck
    In IJCAI 2023 Tutorial, Jul 2023
  9. Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs
    Albert Qiaochu Jiang, Sean Welleck, Jin Peng Zhou, and 6 more authors
    In The Eleventh International Conference on Learning Representations , Jul 2023
    ICLR 2023 Oral
  10. Faith and Fate: Limits of Transformers on Compositionality
    Nouha Dziri, Ximing Lu, Melanie Sclar, and 13 more authors
    In Thirty-seventh Conference on Neural Information Processing Systems, Jul 2023
    NeurIPS 2023 Spotlight

2022

  1. Rainier: Reinforced Knowledge Introspector for Commonsense Question Answering
    Jiacheng Liu, Skyler Hallinan, Ximing Lu, and 4 more authors
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Dec 2022
  2. Peter West, Chandra Bhagavatula, Jack Hessel, and 6 more authors
    In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jul 2022
  3. Jiacheng Liu, Alisa Liu, Ximing Lu, and 5 more authors
    In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), May 2022
  4. Daniel Khashabi, Xinxi Lyu, Sewon Min, and 8 more authors
    In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jul 2022
  5. LILA: A Unified Benchmark for Mathematical Reasoning
    Swaroop Mishra, Matthew Finlayson, Pan Lu, and 8 more authors
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Dec 2022
  6. NaturalProver: Grounded Mathematical Proof Generation with Language Models
    Sean Welleck, Jiacheng Liu, Ximing Lu, and 2 more authors
    In Advances in Neural Information Processing Systems, Dec 2022
  7. COLD Decoding: Energy-based Constrained Text Generation with Langevin Dynamics
    Lianhui Qin, Sean Welleck, Daniel Khashabi, and 1 more author
    In Advances in Neural Information Processing Systems, Dec 2022
    NeurIPS 2022 Oral
  8. QUARK: Controllable Text Generation with Reinforced Unlearning
    Ximing Lu, Sean Welleck, Jack Hessel, and 5 more authors
    In Advances in Neural Information Processing Systems, Dec 2022
    NeurIPS 2022 Oral
  9. NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics
    Ximing Lu, Sean Welleck, Peter West, and 9 more authors
    In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jul 2022
    Best Paper Award
  10. Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations
    Jaehun Jung, Lianhui Qin, Sean Welleck, and 4 more authors
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Dec 2022
  11. Symbolic Brittleness in Sequence Models: on Systematic Generalization in Symbolic Mathematics
    Sean Welleck, Peter West, Jize Cao, and 1 more author
    In AAAI Conference on Artificial Intelligence, Dec 2022

2021

  1. NaturalProofs: Mathematical Theorem Proving in Natural Language
    Sean Welleck, Jiacheng Liu, Ronan Le Bras, and 3 more authors
    In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), Dec 2021
    NeurIPS 2021 Oral
  2. Towards Grounded Natural Language Proof Generation
    Sean Welleck, Jiacheng Liu, Jesse Michael Han, and 1 more author
    In NeurIPS 2021 Workshop on Math AI for Education: Bridging the Gap Between Research and Smart Education, Dec 2021
  3. Divergence Frontiers for Generative Models: Sample Complexity, Quantization Effects, and Frontier Integrals
    Lang Liu, Krishna Pillutla, Sean Welleck, and 3 more authors
    In Advances in Neural Information Processing Systems, Dec 2021
  4. Ilia Kulikov, Sean Welleck, and Kyunghyun Cho
    In Proceedings of the 5th Workshop on Structured Prediction for NLP (SPNLP 2021), Aug 2021
  5. MLE-guided parameter search for task loss minimization in neural sequence modeling
    Sean Welleck, and Kyunghyun Cho
    In AAAI Conference on Artificial Intelligence, Aug 2021
  6. MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
    Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, and 4 more authors
    In Advances in Neural Information Processing Systems, Aug 2021
    NeurIPS 2021 Outstanding Paper Award

2020

  1. Consistency of a Recurrent Language Model With Respect to Incomplete Decoding
    Sean Welleck, Ilia Kulikov, Jaedeok Kim, and 2 more authors
    In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov 2020
  2. Don’t Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training
    Margaret Li, Stephen Roller, Ilia Kulikov, and 4 more authors
    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020
  3. A Generalized Framework of Sequence Generation with Application to Undirected Sequence Models
    Elman Mansimov, Alex Wang, Sean Welleck, and 1 more author
    Jul 2020
  4. Neural Text Generation With Unlikelihood Training
    Sean Welleck, Ilia Kulikov, Stephen Roller, and 3 more authors
    In International Conference on Learning Representations, Jul 2020

2019

  1. Non-Monotonic Sequential Text Generation
    Sean Welleck, Kianté Brantley, Hal Daumé Iii, and 1 more author
    In Proceedings of the 36th International Conference on Machine Learning, 09–15 jun 2019
  2. Sean Welleck, Jason Weston, Arthur Szlam, and 1 more author
    In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul 2019
  3. Sean Welleck, and Kyunghyun Cho
    In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019), Sep 2019

2018

  1. Sean Welleck, Zi-Mei Yao, Yujie Gai, and 3 more authors
    In Neural Information Processing Systems, Sep 2018

2017

  1. Sean Welleck, Jialin Mao, Kyunghyun Cho, and 1 more author
    In Advances in Neural Information Processing Systems, Sep 2017