Evaluation and Generalization

We develop rigorous ways to evaluate language models and understand when, how, and why they generalize.

Related publications

2026

  1. On the Limits and Opportunities of AI Reviewers: Reviewing the Reviews of Nature-Family Papers with 45 Expert Scientists
    Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, and 55 more authors
    arXiv, Dec 2026
  2. Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
    Guijin Son, Seungone Kim, Catherine Arnett, and 73 more authors
    arXiv, Dec 2026
  3. Mind the Sim2Real Gap in User Simulation for Agentic Tasks
    Xuhui Zhou, Weiwei Sun, Qianou Ma, and 8 more authors
    Dec 2026
    COLM 2026
  4. Scaling Evaluation-Time Compute with Reasoning Models as Evaluators
    Seungone Kim, Ian Wu, Jinu Lee, and 8 more authors
    In Findings of the Association for Computational Linguistics: ACL 2026, Dec 2026
    ACL 2026 Findings
  5. RefineBench: Evaluating Refinement Capability of Language Models via Checklists
    Young-Jun Lee, Seungone Kim, Byung-Kwan Lee, and 6 more authors
    Dec 2026
    ICLR 2026
  6. OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
    Pranjal AggarwalSeungone Kim, Jack Lanchantin, and 4 more authors
    Dec 2026
    ICLR 2026

2025

  1. The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
    Seungone Kim, Juyoung Suk, Ji Yong Cho, and 29 more authors
    In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Dec 2025
    Best Paper Award
  2. Evaluating Language Models as Synthetic Data Generators
    Seungone Kim, Juyoung Suk, Xiang Yue, and 7 more authors
    In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Jul 2025

2024

  1. Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
    Zhiqing Sun, Longhui Yu, Yikang Shen, and 4 more authors
    In Advances in Neural Information Processing Systems, Jul 2024
  2. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
    Seungone Kim, Juyoung Suk, Shayne Longpre, and 7 more authors
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Nov 2024

2023

  1. MAUVE Scores for Generative Models: Theory and Practice
    Krishna Pillutla, Lang Liu, John Thickstun, and 6 more authors
    JMLR, Nov 2023
  2. Faith and Fate: Limits of Transformers on Compositionality
    Nouha Dziri, Ximing Lu, Melanie Sclar, and 13 more authors
    In Thirty-seventh Conference on Neural Information Processing Systems, Nov 2023
    NeurIPS 2023 Spotlight

2022

  1. Symbolic Brittleness in Sequence Models: on Systematic Generalization in Symbolic Mathematics
    Sean Welleck, Peter West, Jize Cao, and 1 more author
    In AAAI Conference on Artificial Intelligence, Nov 2022

2021

  1. Divergence Frontiers for Generative Models: Sample Complexity, Quantization Effects, and Frontier Integrals
    Lang Liu, Krishna Pillutla, Sean Welleck, and 3 more authors
    In Advances in Neural Information Processing Systems, Nov 2021
  2. MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
    Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, and 4 more authors
    In Advances in Neural Information Processing Systems, Nov 2021
    NeurIPS 2021 Outstanding Paper Award