Learning and Optimization

We study how objectives, data, and optimization shape what language models learn and how effectively they generalize.

Related publications

2026

  1. GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
    Ningyuan Yang, Weihua Du, Weiwei Sun, and 2 more authors
    Nov 2026
    COLM 2026

2025

  1. Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
    Andre He, Daniel Fried, and Sean Welleck
    Nov 2025
    EMNLP 2025 Oral
  2. Agentic-R1: Distilled Dual-Strategy Reasoning
    Weihua DuPranjal AggarwalSean Welleck, and 1 more author
    arXiv, Nov 2025
    EMNLP 2025
  3. L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
    Pranjal Aggarwal, and Sean Welleck
    Nov 2025
    COLM 2025

2024

  1. Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
    Zhiqing Sun, Longhui Yu, Yikang Shen, and 4 more authors
    In Advances in Neural Information Processing Systems, Nov 2024

2023

  1. Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning
    Ximing Lu, Faeze Brahman, Peter West, and 14 more authors
    EMNLP, Nov 2023
  2. STEER: Unified Style Transfer with Expert Reinforcement
    Skyler Hallinan, Faeze Brahman, Ximing Lu, and 3 more authors
    EMNLP Findings, Nov 2023
  3. Generating Sequences by Learning to Self-Correct
    Sean Welleck, Ximing Lu, Peter West, and 4 more authors
    In The Eleventh International Conference on Learning Representations , Nov 2023

2022

  1. Rainier: Reinforced Knowledge Introspector for Commonsense Question Answering
    Jiacheng Liu, Skyler Hallinan, Ximing Lu, and 4 more authors
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Dec 2022
  2. QUARK: Controllable Text Generation with Reinforced Unlearning
    Ximing Lu, Sean Welleck, Jack Hessel, and 5 more authors
    In Advances in Neural Information Processing Systems, Dec 2022
    NeurIPS 2022 Oral

2021

  1. MLE-guided parameter search for task loss minimization in neural sequence modeling
    Sean Welleck, and Kyunghyun Cho
    In AAAI Conference on Artificial Intelligence, Dec 2021

2020

  1. Don’t Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training
    Margaret Li, Stephen Roller, Ilia Kulikov, and 4 more authors
    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul 2020
  2. Neural Text Generation With Unlikelihood Training
    Sean Welleck, Ilia Kulikov, Stephen Roller, and 3 more authors
    In International Conference on Learning Representations, Jul 2020