Reasoning Models

We aim to make reasoning models more capable, controllable, and efficient, including through inference-time scaling and efficient inference.

Related publications

2026

  1. Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
    Guijin Son, Seungone Kim, Catherine Arnett, and 73 more authors
    arXiv, 09–15 jun 2026
  2. Reasoning over Mathematical Objects: On-Policy Reward Modeling and Test Time Aggregation
    Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim, and 18 more authors
    arXiv, 09–15 jun 2026
  3. Argument Reconstruction as Supervision for Critical Thinking in LLMs
    Hyun Ryu, Gyouk Chu, Gregor Betz, and 3 more authors
    arXiv, 09–15 jun 2026
  4. Scaling Evaluation-Time Compute with Reasoning Models as Evaluators
    Seungone Kim, Ian Wu, Jinu Lee, and 8 more authors
    In Findings of the Association for Computational Linguistics: ACL 2026, 09–15 jun 2026
    ACL 2026 Findings
  5. OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
    Pranjal AggarwalSeungone Kim, Jack Lanchantin, and 4 more authors
    09–15 jun 2026
    ICLR 2026
  6. The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
    Seongyun Lee, Seungone Kim, Minju Seo, and 9 more authors
    09–15 jun 2026
    ICLR 2026

2025

  1. Agentic-R1: Distilled Dual-Strategy Reasoning
    Weihua DuPranjal AggarwalSean Welleck, and 1 more author
    arXiv, 09–15 jun 2025
    EMNLP 2025
  2. L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
    Pranjal Aggarwal, and Sean Welleck
    09–15 jun 2025
    COLM 2025

2023

  1. A Survey of Deep Learning for Mathematical Reasoning
    Pan Lu, Liang Qiu, Wenhao Yu, and 2 more authors
    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2023

2022

  1. LILA: A Unified Benchmark for Mathematical Reasoning
    Swaroop Mishra, Matthew Finlayson, Pan Lu, and 8 more authors
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Dec 2022
  2. NaturalProver: Grounded Mathematical Proof Generation with Language Models
    Sean Welleck, Jiacheng Liu, Ximing Lu, and 2 more authors
    In Advances in Neural Information Processing Systems, Dec 2022
  3. Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations
    Jaehun Jung, Lianhui Qin, Sean Welleck, and 4 more authors
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Dec 2022

2021

  1. NaturalProofs: Mathematical Theorem Proving in Natural Language
    Sean Welleck, Jiacheng Liu, Ronan Le Bras, and 3 more authors
    In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), Dec 2021
    NeurIPS 2021 Oral
  2. Towards Grounded Natural Language Proof Generation
    Sean Welleck, Jiacheng Liu, Jesse Michael Han, and 1 more author
    In NeurIPS 2021 Workshop on Math AI for Education: Bridging the Gap Between Research and Smart Education, Dec 2021