Teng Xiao
Teng Xiao

Teng Xiao

I work at the Allen Institute for AI (AI2) and the University of Washington with Prof. Noah A. Smith and Prof. Hanna Hajishirzi. I completed my PhD at Pennsylvania State University, advised by Prof. Vasant Honavar.

I am interested in machine learning and reinforcement learning. I am currently working on: (i) Alignment and Reasoning for LLM; (ii) Long-horizon LLM Agent for Open-ended Tasks. My work is regularly published at ICML, ICLR, and NeurIPS.

News

Selected Research

A selection organised by theme; the full list is on Google Scholar. * equal contribution, advising role.

01

Self-Evolution for LLMs

  1. 2026
    Rethinking the Evaluation of Harness Evolution for Agents

    Yike Wang*, Huaisheng Zhu*, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao*

    Preprint Paper Blog Code

  2. 2026
    Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

    Zhengyu Chen*, Teng Xiao*, Huaisheng Zhu, Yige Yuan, Luan Zhang, Jingang Wang

    Preprint Paper

  3. 2026
    Meta-Reinforcement Learning with Self-Reflection for Agentic Search

    Teng Xiao, Yige Yuan*, Hamish Ivison, Huaisheng Zhu, Faeze Brahman, Nathan Lambert, Pradeep Dasigi, Noah A. Smith, Hannaneh Hajishirzi

    COLM Paper Code

02

Reinforcement Learning for LLMs

  1. 2026
    Tmax: A Simple Recipe for Terminal Agents

    Hamish Ivison*, Junjie Oscar Yin*, Rulin Shao, Teng Xiao, Nathan Lambert, Hannaneh Hajishirzi

    Preprint Paper Code

  2. 2025
    Olmo 3

    OLMo Team, including Teng Xiao*

    Core contributor, RL post-training.

    Tech Report Paper Models

  3. 2025
    Internalizing World Models via Self-Play Finetuning for Agentic RL

    Shiqi Chen*, Tongyao Zhu*, Zian Wang*, Jinghan Zhang*, Kangrui Wang, Siyang Gao, Teng Xiao*, Yee Whye Teh, Junxian He, Manling Li

    Preprint Paper Code

  4. 2025
    Inference-time Alignment in Continuous Space

    Yige Yuan, Teng Xiao*, Yunfan Li, Bingbing Xu, Shuchang Tao, Yunqi Qiu, Huawei Shen, Xueqi Cheng

    NeurIPS Paper

  5. 2025
    SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters

    Teng Xiao, Yige Yuan*, Zhengyu Chen, Mingxiao Li, Shangsong Liang, Zhaochun Ren, Vasant G. Honavar

    Used by the EXAONE series models from LG AI Research.

    ICLR Paper Code

  6. 2025
    On a Connection Between Imitation Learning and RLHF

    Teng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen, Vasant G. Honavar

    ICLR Paper Code

  7. 2024
    Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

    Teng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li, Vasant G. Honavar

    NeurIPS Paper Code

03

Structured Knowledge Representation Learning

  1. 2024
  2. 2024
    Efficient Contrastive Learning for Fast and Accurate Inference on Graphs

    Teng Xiao, Huaisheng Zhu, Zhiwei Zhang, Zhimeng Guo, Charu C. Aggarwal, Suhang Wang, Vasant G. Honavar

    ICML Paper Code

  3. 2023
    Simple and Asymmetric Graph Contrastive Learning without Augmentations

    Teng Xiao*, Huaisheng Zhu*, Zhengyu Chen, Suhang Wang

    NeurIPS Paper Code

  4. 2022
    Decoupled Self-supervised Learning for Graphs Spotlight · top 5%

    Teng Xiao, Zhengyu Chen, Zhimeng Guo, Zeyang Zhuang, Suhang Wang

    NeurIPS Paper Code

Academic Service

Area Chair
NeurIPS, ACL
Program Committee & Reviewer
ICLR (2022–2026), NeurIPS (2022–2025), ICML (2023, 2024), WSDM (2023–2025), AAAI (2022, 2023), SIGIR (2021–2023), ACL ARR (2024), CIKM (2023), TheWebConf (2022, 2023), COLM (2024)
Journal Reviewer
ACM Transactions on Intelligent Systems and Technology, ACM Transactions on Information Systems